OpenAI Released Automated AI Security Testing Results
The new GPT-Red system outperformed human red-teamers at finding indirect prompt-injection vulnerabilities.
Updated on Sept. 26, 2026 in Artificial Intelligence

Live Poll
Do you trust current automated security testing to prevent malicious attacks on AI tools you use?
OpenAI disclosed results for GPT-Red, an automated security testing system that identified weaknesses in AI models. Researchers found that GPT-Red succeeded in 84% of indirect prompt-injection scenarios, significantly outpacing the 13% success rate of human red-teamers.
Why it matters
Automated testing systems like GPT-Red are designed to identify system vulnerabilities to train production models to resist failures. This shift toward autonomous red-teaming addresses the escalating difficulty of manual security assessments in increasingly complex large language models.
GPT-Red achieved an 84% success rate in indirect prompt-injection tests, compared to 13% for human testers. Additionally, GPT-5.6 Sol demonstrated a 0.05% failure rate in latest evaluations, reflecting a six-fold reduction in failures compared to its predecessor.
The players
OpenAI
An AI research and deployment company focused on the development of large language models and automated safety testing systems.
The details
GPT-Red utilizes self-play reinforcement learning, a method where the system iterates against itself to identify and exploit vulnerabilities. Separately, Morris-II research employs retrieval-augmented generation—a technique that connects AI models to external data sources—to pass malicious instructions between email assistants. To defend against such threats, the Virtual Donkey defense recorded a 1.0 true-positive rate and a 0.015 false-positive rate.
Timeline
May 2026: Best production model benchmarks were established.
September 26, 2026: OpenAI disclosed the findings for the GPT-Red system.
The Tech Race
This development follows the pattern set by the Morris-II research, which first demonstrated self-replicating prompts between simulated AI assistants. The race now focuses on scaling automated defenses, such as the Virtual Donkey, to match the speed of autonomous red-teaming systems.
These improvements in model security will eventually result in more resilient AI assistants for enterprise and consumer users. Users should watch for updates to model safety protocols as vendors integrate these automated testing systems into their release cycles.
The takeaway
The move toward automated red-teaming signals a fundamental shift in how AI vulnerabilities are identified and mitigated before release. Researchers should track the future integration of the Virtual Donkey defense and similar protocols in upcoming model deployments.
Further reading
For more on how language models are tested for safety, see our latest coverage on Artificial Intelligence.
Source note: This article includes information reported by TokenPost.
Live Poll
Do you trust current automated security testing to prevent malicious attacks on AI tools you use?








