Researchers introduced SWE-bench Science, a repository-level benchmark comprising 119 tasks drawn from 98 GitHub repositories across 20 scientific domains, to test whether coding agents can repair scientific software without compromising scientific validity. Even the top-performing agent achieved below a 50% success rate, with failures traced to four recurring patterns: gaps in scientific knowledge, superficial repairs, incomplete coverage, and poor generalization. The team also found that providing scientific guidance produced mixed effects, in some cases constraining rather than improving agent performance.