JIT-Agent: Just-in-Time Synthesis of Agent Harnesses
JIT-Agent is a harness intelligence model that generates task-adaptive agent harnesses on the fly for arbitrary off-the-shelf LLMs, formalizing the…
JIT-Agent is a harness intelligence model that generates task-adaptive agent harnesses on the fly for arbitrary off-the-shelf LLMs, formalizing the…
D3-MOPD is a zero-overhead scheduler for multi-teacher on-policy distillation that adjusts per-domain training mixture ratios online, reusing the reverse-KL signal…
Agent-G2 addresses reward sparsity in long-horizon agentic RL by replacing the common practice of treating expert-trajectory guidance depth as a…
The Christian Science Monitor’s book critics sampled eleven nonfiction and fiction titles addressing artificial intelligence to help readers understand the…
Code World Model separates world-state evolution from visual rendering by using a coding agent to reason about events and consequences…
This paper runs a controlled comparison of next-chunk reasoning reinforcement learning against a simple alternative called Mixed SFT — a…
This paper identifies a ‘capability integration gap’ in multi-teacher on-policy distillation, where a controlled benchmark shows standard methods recover only…
SWE Refactor Bench is a new benchmark of 20 whole-repository code migrations across four categories of technical debt, designed to…
AnTrap is a benchmark that injects realistic dynamic anomalies — such as unexpected pop-ups and action misuse — into Android…
This study measures the cost-quality trade-off of switching between cheaper and more capable models partway through a long-running coding agent…
Gated Recurrent Transformer is a new architecture that brackets a single shared transformer core, iterated multiple times, between fixed prelude…
Prefix Sliding is an inference technique that discards reasoning tokens outside of a fixed prefix (containing key instructions and tools)…