A new study finds that real-world coding requests differ sharply from the curated GitHub issues used in SWE-bench-style benchmarks: 88% of real prompts contain only a bare problem statement versus just 7% of benchmark tasks, and are far more casually written. The authors introduce RealSWE, 381 task-family variants derived from SWE-bench Verified and Pro that vary information composition and linguistic style, and find that realistic inputs reduce coding-agent resolution rates by 6.4 percentage points on average and can reorder model rankings, with explicitly stating desired behavior and motivation improving performance the most.
