Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
A study tested whether committees of LLM agents used for clinical decision support can be manipulated through benchmark shortcuts that…
A study tested whether committees of LLM agents used for clinical decision support can be manipulated through benchmark shortcuts that…
Researchers proposed S²VOPD, a self-supervised on-policy distillation method that creates teacher-student asymmetry by degrading the student’s input rather than augmenting…
SimpleOPD is a new on-policy distillation method for transferring long-context reasoning ability from large teacher models to smaller, shorter-context student…
A paper introduces LOPD, a method that makes an agent’s own privileged context end-to-end learnable from experience rather than relying…
Researchers from Peking University and BUPT identified a failure mode in sequential reinforcement learning with verifiable rewards, where optimizing for…
WeiboAI researchers introduced CLR (Claim-Level Reliability Assessment), a training-free framework that improves LLM reasoning accuracy by directing verification effort at…
Researchers released PRM-as-a-Judge 1.5, an updated toolkit that evaluates robotic manipulation policies by converting video rollouts into fine-grained progress metrics…
Researchers at the University of Milan-Bicocca built ParliamentRAG, a retrieval-augmented generation system for Italian Chamber of Deputies records that scores…
A new benchmark called MobileMem evaluates AI agents’ ability to provide persistent personal assistance by synthesizing a full year of…
Researchers introduced Mobius-v0, an architecture that separates knowledge storage from reasoning by using a globally shared memory implemented as a…
Researchers released Mimir v1, a 1-billion-parameter language model built on the Hierarchical Reasoning Model architecture and trained exclusively on ethically-sourced,…
A study examines how many times high-quality domain-specific data should be repeated during LLM pretraining as models and token budgets…