Training Agents to Evolve with Their Harness: The TaoLive Digital Avatar System
A technical report from Alibaba’s Taobao Live team describes Harness-Aware Training (HAT), a method for training compact models to adapt…
A technical report from Alibaba’s Taobao Live team describes Harness-Aware Training (HAT), a method for training compact models to adapt…
This paper introduces PILOT, a supervisor-worker agent harness that performs self-improvement live during a run rather than only after it…
This study systematically compares Evolution Strategies (ES) against Group Relative Policy Optimization (GRPO) as post-training paradigms for LLM reasoning. It…
WikiSkill is a framework that co-evolves an LLM agent’s reusable skill library alongside a persistent wiki-style knowledge base, separating raw…
CaSKG is a retrieval framework that calibrates the reliability of a skill graph’s edges before using it to retrieve procedural…
This paper formalizes what is required to legitimately ‘replay’ a claim attached to an LLM evaluation metric, then applies that…
VoiceMem is a memory architecture for duplex speech language models built around a parallel ‘informational left brain’ and ’emotional right…
FrontierChallenge is a new cross-domain benchmark of 300 end-to-end scientific workflows spanning quantum chemistry, molecular dynamics, materials science, and other…
This paper shows that common off-policy RL stabilizers behave differently depending on data regime: parameter normalization helps under narrow replay…
JIT-Agent is a harness intelligence model that generates task-adaptive agent harnesses on the fly for arbitrary off-the-shelf LLMs, formalizing the…
D3-MOPD is a zero-overhead scheduler for multi-teacher on-policy distillation that adjusts per-domain training mixture ratios online, reusing the reverse-KL signal…
Agent-G2 addresses reward sparsity in long-horizon agentic RL by replacing the common practice of treating expert-trajectory guidance depth as a…