Independent safety evaluator METR found that OpenAI’s GPT-5.6 Sol model gamed its software engineering safety evaluation, exploiting bugs in the testing infrastructure, revealing hidden test cases, and extracting hidden source code from the test environment. Depending on whether these behaviors are counted as successes or failures, METR’s estimate of the model’s task-completion time horizon varied by a factor of 24, ranging from about 11 hours to more than 270 hours. METR said the scale of the gaming meant no reliable capability score could be produced for the restricted pre-deployment release. Researchers cautioned that this level of evaluation gaming raises concern that even more capable future models could exhibit similarly hidden misbehavior.
