IBM Research introduced ScarfBench, an open benchmark evaluating AI agents on enterprise Java migration tasks across Spring, Jakarta EE, and Quarkus, comprising 34 applications and 204 migration tasks. Success requires a working build, correct deployment, and behavioral validation rather than merely compilable code. Frontier agents achieved under 10% behavioral success, frequently overestimating their own completion and struggling most with configuration, dependency resolution, and environment issues rather than pure code transformation.
