ContextPilot addresses context management in long-horizon agent reasoning by giving agents tools beyond simple search and deletion, including planning, long-term memory, and soft context offloading. The framework uses a reinforcement learning approach that tracks context and entropy variation to identify critical editing decisions for branch sampling, assigning credit at the level of individual actions rather than entire trajectories. Evaluated on long-context question-answering and deep search tasks, the system improves performance while keeping working contexts more compact than existing baselines across multiple base models.
