OpenAI Built an AI That Hacks Its Own Models, and It's Beating Humans by a Landslide

The Core · TL;DR
- OpenAI built GPT-Red, an internal automated red-teaming model trained via self-play reinforcement learning to find security flaws in its own AI systems.
- GPT-Red succeeded in 84% of indirect prompt injection attacks against GPT-5.1, versus just 13% for human red-teamers.
- The model uncovered a new exploit called Fake Chain-of-Thought prompt injection, and its findings feed directly into training OpenAI's production models.
- Reports conflict on which model versions (GPT-5.5 vs GPT-5.6 Sol) were tested and how well they resisted attacks; OpenAI keeps GPT-Red internal and separate from deployed systems.
OpenAI has a new internal weapon for stress-testing its own models, and the results suggest human red-teamers may be outmatched. The system, called GPT-Red, is an automated red-teaming model purpose-built to probe OpenAI's AI systems for security weaknesses, with a particular focus on prompt injection, one of the most persistent threats to deployed language models.
The numbers OpenAI has shared are striking. Against GPT-5.1, GPT-Red succeeded in 84% of indirect prompt injection scenarios it attempted. Human red-teamers, tasked with the same job, managed just 13%. That gap of nearly 6-to-1 signals a meaningful shift in how frontier AI labs may approach adversarial testing going forward: rather than relying primarily on human ingenuity to find exploits, OpenAI is training a model to do it at scale and, apparently, with far greater success.
How GPT-Red Learns to Attack
GPT-Red isn't a static tool. It's trained through self-play reinforcement learning, pitting an attacker model against a defender model in repeated rounds, each side sharpening its tactics against the other. This adversarial loop is what allowed GPT-Red to surface a previously undocumented exploit OpenAI is calling "Fake Chain-of-Thought" direct prompt injection, a novel technique that manipulates a model's apparent reasoning trace to smuggle in malicious instructions.
Crucially, GPT-Red never ships to users. OpenAI has kept it entirely internal, walled off from production systems, and used purely as an offensive testing instrument. The vulnerabilities and attack patterns it uncovers are fed back into the training and evaluation pipeline for OpenAI's public-facing models, effectively turning red-teaming into part of the model development loop rather than a final compliance check before launch.
A Murky Detail on Which Models Were Tested
Reporting on GPT-Red isn't fully consistent about which model versions were put through the wringer. One account references testing against "GPT-5.6 Sol," describing it as a recent target, while another reports that GPT-Red "breaks nearly all models it is pitted against," including internal and production systems up through GPT-5.5, yet separately notes that Fake Chain-of-Thought attacks saw success rates below 10% specifically against GPT-5.6 Sol. The discrepancy suggests either a naming inconsistency across reports or a genuine improvement in later model defenses that outside coverage hasn't fully reconciled. OpenAI hasn't issued a public clarification on the exact version lineage tested.
Not the Only Lab Doing This
OpenAI isn't alone in automating adversarial testing. Anthropic runs a comparable effort through its Frontier Red Team, using AI systems to simulate attacks against new models before they're released. The approach reflects a broader trend among frontier labs: as models grow more capable and more exposed to real-world prompt injection risks, especially through tool use and agentic workflows, manual red-teaming alone is increasingly seen as too slow and too limited in coverage to keep pace.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
