Researchers introduced AgentMercury, a framework that synthesizes persistent, executable business environments — complete with entities, services, and cross-service rules — rather than relying on predefined, task-centric benchmarks. The system generated 4,783 environments spanning 14 industries and 50 countries, from which diverse enterprise tasks emerge naturally. Training agents on these scenario-grounded environments improved performance on enterprise workflows and out-of-domain benchmarks, with one model’s EnterpriseOps-GYM score rising from 12.3 to 15.7, and the authors show the environment-construction process itself can be learned as a capability.
