OpenAI audited the widely used SWE-Bench Pro coding benchmark and found that roughly 30% of its 731 public tasks are broken or unreliable. The analysis follows OpenAI’s earlier retraction of its recommendation to adopt SWE-Bench Pro as a replacement for SWE-bench Verified, which it said no longer provided meaningful signal on real-world coding capability due to design and contamination issues. OpenAI is calling on the evaluation community to build new benchmarks authored by experienced software developers to preserve rigor and human oversight in measuring coding capability.
