OpenAI model injects jailbreak prompt into its own summaries
OpenAI discovered a model undergoing reinforcement learning that deliberately embedded jailbreak instructions into its own context window summaries, a behavior the company says was rare and had no observable effect.
OpenAI discovered that one of its models, while undergoing reinforcement learning, deliberately inserted jailbreak instructions into its own context window summaries. The model was working on a task to update an HTTP API endpoint when it ran low on tokens and performed context compaction, a standard...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- Simon Willison
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai