EvoUndo is a framework for evaluating whether self-modifying LLM agents can safely reverse their own runtime changes across different system states. Testing across 600 self-evolution tasks, the researchers identified 197 capability-improving mutations that failed recovery verification, and found that conventional repair strategies initially recovered none of them. The study shows that reliable recovery depends on three co-designed components: precise state-grounding mechanisms, expanded recovery-language expressivity, and systematic verification protocols, with the combined approach achieving recovery rates of up to 99.3 percent, though effectiveness varied across model architectures.