SWE Refactor Bench is a new benchmark of 20 whole-repository code migrations across four categories of technical debt, designed to catch agents that pass tests by quietly copying the original implementation instead of performing the migration — a failure mode the authors call ‘Blindness.’ Its three-stage evaluation protocol checks that a migration actually occurred, that behavior is preserved, and runs independent agent-generated tests for hidden differences. Across 520 runs from 8 frontier models, only 28 runs (5.4%) passed all three stages, and the best model, Claude Opus 5, scored just 47.0 out of 100.