A new paper proposes environment evolution, a method that incrementally increases the difficulty of training environments for terminal agents off-policy across training generations, addressing how synthesized environments quickly become too easy for increasingly capable frontier models. Implemented as a loop-engineered multi-agent harness and validated with Claude Opus 5, GPT-5.6 Sol, and Hy4 preview, the technique consistently produces harder environments and, when used for long-horizon RL training on Qwen3.6-27B and Qwen3.6-35B-A3B, improves their Terminal-Bench 2.1 performance by 14.4 and 18.0 percentage points respectively.