OpenAI developed GPT-Red, an automated red-teaming model trained through self-play reinforcement learning to identify prompt injection vulnerabilities in AI systems. The system achieved an 84% success rate on indirect prompt injection tests, outperforming human red-teamers on the same scenarios. GPT-Red’s attack methods are fed back into production model training, enabling GPT-5.6 Sol to reduce direct prompt injection failures to 0.05% while maintaining consistent capability performance.