Neural Edition
Artificial Intelligence
OpenAI GPT-Red Beats Human Red-Teamers 84% to 13% on Prompt Injection
OpenAI just showed an internal attacker model, GPT-Red, that beat human red-teamers 84% to 13% on a held-out prompt-injection arena—and used those attacks to harden GPT-5.6 Sol.
What are we talking about?
- Prompt injection — sneaky instructions planted in text an AI reads (a webpage, email, file, or tool reply) that try to make the model ignore its real job.
- Red-teaming — deliberate attacking of a system to find weak spots before bad actors do.
- GPT-Red — OpenAI’s automated red-teaming model. It sends an attack, watches the reply, and tries again—like a human tester, but at machine scale.
- Self-play — GPT-Red attacks a pool of defender models while those defenders train to resist and still finish their tasks.
- GPT-5.6 Sol — OpenAI’s latest production model trained against GPT-Red; the company calls it its most robust release yet against prompt injection.
How it works
GPT-Red crafts an injection
→
Defender GPT replies
→
Reward: attack success vs task done
→
Stronger attacks enter training
- OpenAI places GPT-Red in simulated “dojos”: browsing, email/calendar, code editing, and tool use—places where third-party text can hide instructions.
- Each environment defines what GPT-Red controls (a file snippet, a banner, an email body, tool output) and what counts as a valid failure.
- GPT-Red is rewarded for eliciting a real failure; defenders are rewarded for resisting without refusing everything.
- As defenders harden, GPT-Red is forced into stronger, more diverse attacks—including a novel class OpenAI calls fake chain-of-thought, where a spoofed reasoning note tricks the target.
- Those attacks feed the training of production models. OpenAI says it has used this loop since GPT-5.3 through GPT-5.6 Sol.
What changed
Before GPT-Red scale-up
Human red-teaming is slow. On the replicated Dziemian et al. (2025) indirect prompt-injection arena against GPT-5.1, humans succeeded on 13% of scenarios. OpenAI says common robustness evals were already saturated.
After folding GPT-Red into training
On the same arena, GPT-Red succeeded on 84% of scenarios. GPT-5.6 Sol shows 6× fewer failures on OpenAI’s hardest direct prompt-injection benchmark than its best production model from four months earlier; held-out GPT-Red direct injections fail on 0.05% of cases. Strongest GPT-Red attacks: over 90% hit GPT-5 (Aug 2025) vs under 23% on GPT-5.6.
84% vs 13%
GPT-Red vs human success on the replicated indirect prompt-injection arena (target: GPT-5.1)
Why it matters
- Agents raise the blast radius. Models that browse, read files, and call tools can encounter attacker-controlled text outside the chat box.
- Safety can get a flywheel. OpenAI frames GPT-Red as the start of using today’s models to harden tomorrow’s—mirroring how AI already helps train next-gen capabilities.
- Novel attacks surface earlier. Fake chain-of-thought was new to OpenAI’s researchers; on that class, success fell from upwards of 95% on GPT-5.1 to below 10% on GPT-5.6 Sol.
- Live agents are in scope. Against Andon Labs’ office vending agent “Vendy,” GPT-Red cut an item to the $0.50 floor, listed a $100+ item at $0.50, and canceled another customer’s order—then OpenAI disclosed the flaws.
What it does NOT prove / caveats
- Company-reported benchmarks. The 84%/13%, 6×, 0.05%, and >97% indirect-benchmark figures come from OpenAI’s announcement and coverage summarizing it—not an independent audit.
- GPT-Red stays internal. OpenAI will not release it; the offensive capability is kept off deployed products.
- Humans still matter. OpenAI and outside analysts note GPT-Red is weaker on multi-turn chats and image-based injections; people still find attacks it misses.
- Robust ≠ invincible. Lower failure rates on OpenAI’s suites do not prove immunity to every real-world harness or future attack class.
- Pre-print pending. OpenAI said more technical detail would follow in a pre-print “later this week” from the July 15, 2026 post.
End-to-end recap
Five steps
- Agents read third-party text → prompt injection risk grows.
- GPT-Red self-plays as an attacker against defender GPTs.
- It beats humans 84% to 13% on a held-out injection arena and invents fake chain-of-thought.
- Those attacks train GPT-5.6 Sol → 6× fewer failures on the hardest direct test.
- Live case study: Vendy’s $0.50 price hacks get disclosed; GPT-Red itself stays locked inside OpenAI.
Editor’s note: Neural Edition summarizes public reporting and labels company or founder claims as such. How we report · Corrections
