Anthropic researchers discovered that infrastructure configuration significantly impacts agentic coding benchmarks, with resource allocation alone causing score variations of up to 6 percentage points on Terminal-Bench 2.0 — sometimes exceeding the gaps between top-ranked models. Container runtime resource enforcement operates via dual parameters — guaranteed allocation and hard kill threshold — and when these are set identically, transient memory spikes trigger unexpected failures unrelated to model capability. The findings highlight a critical blind spot in current AI evaluation methodology.
