ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot addresses context management in long-horizon agent reasoning by giving agents tools beyond simple search and deletion, including planning, long-term…
ContextPilot addresses context management in long-horizon agent reasoning by giving agents tools beyond simple search and deletion, including planning, long-term…
EvoUndo is a framework for evaluating whether self-modifying LLM agents can safely reverse their own runtime changes across different system…
StarHarness is a framework that optimizes the executable environment surrounding a fixed language model by evolving prompts, tool interfaces, skills,…
DART-SD addresses limitations in training multi-turn tool-calling agents by modeling task execution as an Interaction-State Transition Graph that captures the…
LoopArena is a benchmark that evaluates how effectively one AI model can direct a separate coding agent through extended, multi-round…
VLANeXt is a research-oriented codebase that systematically explores design choices for vision-language-action models across more than 500 experiments spanning foundational…
A new paper proposes a unified framework for generating high-quality training data for LLM agents, representing agentic data as a…
New research addresses how to improve language model reasoning during inference without repeated generation or external verification systems. The proposed…
GameWAM addresses a gap between game-playing agents, which typically lack explicit modeling of world dynamics, and game world models, which…
A new benchmark investigates whether current video generation models can function as authentic world models by accurately reproducing the distribution…
New research investigates whether multimodal large language models can translate local visual perception into effective spatial navigation within complex urban…
Researchers propose Test-Time Policy Optimization (TTPO), a label-free training objective for adapting language models to math reasoning tasks at test…