The paper introduces OpenART, an open-ended arena for scalable agent red teaming that evolves adversarial test environments and comprises over 10,000 validated scenarios spanning 50 domains. It proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box method that uncovers agent vulnerabilities by modifying environmental states rather than model parameters. Tested across 75 agent configurations, EMHA achieved an 85% attack success rate, with effectiveness increasing substantially as task complexity rises, indicating that implementation details play a significant role in agent safety outcomes.
