TTPO: Test-Time Policy Optimization for Label-Free Reasoning
Researchers propose Test-Time Policy Optimization (TTPO), a label-free training objective for adapting language models to math reasoning tasks at test…
Researchers propose Test-Time Policy Optimization (TTPO), a label-free training objective for adapting language models to math reasoning tasks at test…
A technical report from Alibaba’s Taobao Live team describes Harness-Aware Training (HAT), a method for training compact models to adapt…
This paper introduces PILOT, a supervisor-worker agent harness that performs self-improvement live during a run rather than only after it…
This study systematically compares Evolution Strategies (ES) against Group Relative Policy Optimization (GRPO) as post-training paradigms for LLM reasoning. It…
WikiSkill is a framework that co-evolves an LLM agent’s reusable skill library alongside a persistent wiki-style knowledge base, separating raw…
CaSKG is a retrieval framework that calibrates the reliability of a skill graph’s edges before using it to retrieve procedural…
This paper formalizes what is required to legitimately ‘replay’ a claim attached to an LLM evaluation metric, then applies that…
VoiceMem is a memory architecture for duplex speech language models built around a parallel ‘informational left brain’ and ’emotional right…
Anton Leicht argues that policymakers face a consequential choice between two futures for AI and labor: augmentation, where AI enhances…
Saxon Zvina argues that developing nations in the Global South hold unprecedented strategic leverage in shaping the future of artificial…
FrontierChallenge is a new cross-domain benchmark of 300 end-to-end scientific workflows spanning quantum chemistry, molecular dynamics, materials science, and other…
This paper shows that common off-policy RL stabilizers behave differently depending on data regime: parameter normalization helps under narrow replay…