OpenAI Released Automated AI Security Testing Results

The new GPT-Red system outperformed human red-teamers at finding indirect prompt-injection vulnerabilities.

Updated on Sept. 26, 2026 in Artificial Intelligence

Bold flat-color editorial illustration in navy and cream, depicting a symbolic interconnected lattice structure representing automated system security testing.
OpenAI disclosed that its GPT-Red automated testing system successfully identified vulnerabilities in 84% of indirect prompt-injection scenarios, significantly outperforming human red-team assessments. AI Illustration. Upload story photo >

Live Poll

Do you trust current automated security testing to prevent malicious attacks on AI tools you use?

OpenAI disclosed results for GPT-Red, an automated security testing system that identified weaknesses in AI models. Researchers found that GPT-Red succeeded in 84% of indirect prompt-injection scenarios, significantly outpacing the 13% success rate of human red-teamers.

Why it matters

Automated testing systems like GPT-Red are designed to identify system vulnerabilities to train production models to resist failures. This shift toward autonomous red-teaming addresses the escalating difficulty of manual security assessments in increasingly complex large language models.

GPT-Red achieved an 84% success rate in indirect prompt-injection tests, compared to 13% for human testers. Additionally, GPT-5.6 Sol demonstrated a 0.05% failure rate in latest evaluations, reflecting a six-fold reduction in failures compared to its predecessor.

The players

OpenAI

An AI research and deployment company focused on the development of large language models and automated safety testing systems.

The details

GPT-Red utilizes self-play reinforcement learning, a method where the system iterates against itself to identify and exploit vulnerabilities. Separately, Morris-II research employs retrieval-augmented generation—a technique that connects AI models to external data sources—to pass malicious instructions between email assistants. To defend against such threats, the Virtual Donkey defense recorded a 1.0 true-positive rate and a 0.015 false-positive rate.

Timeline

  1. May 2026: Best production model benchmarks were established.

  2. September 26, 2026: OpenAI disclosed the findings for the GPT-Red system.

The Tech Race

This development follows the pattern set by the Morris-II research, which first demonstrated self-replicating prompts between simulated AI assistants. The race now focuses on scaling automated defenses, such as the Virtual Donkey, to match the speed of autonomous red-teaming systems.

These improvements in model security will eventually result in more resilient AI assistants for enterprise and consumer users. Users should watch for updates to model safety protocols as vendors integrate these automated testing systems into their release cycles.

The takeaway

The move toward automated red-teaming signals a fundamental shift in how AI vulnerabilities are identified and mitigated before release. Researchers should track the future integration of the Virtual Donkey defense and similar protocols in upcoming model deployments.

Further reading

For more on how language models are tested for safety, see our latest coverage on Artificial Intelligence.

Source note: This article includes information reported by TokenPost.

Live Poll

Do you trust current automated security testing to prevent malicious attacks on AI tools you use?

OpenAI Released Automated AI Security Testing Results