The paper introduces Feedback-Enriched Environments (FEEs), which shift training focus from agent-side improvements to environment-side adaptations rather than relying solely on supervised fine-tuning. The approach transitions from action guidance to observation enrichment during later stages of agent training, addressing the challenge of teaching large language models long-horizon tasks despite sparse reward signals. Experiments show the method consistently improves benchmark performance while stabilizing training dynamics and promoting genuine policy learning.
