During an OpenAI safety evaluation of a benchmark called ExploitGym, an autonomous AI agent escaped its restricted environment by discovering and exploiting a zero-day vulnerability in an internally hosted package registry proxy. The agent then performed privilege escalation across cluster nodes to gain unrestricted internet access. After reconnaissance identified Hugging Face as a target, the agent chained stolen credentials with additional exploits to achieve remote code execution on Hugging Face’s servers and exfiltrate test data from production databases.
