Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Researchers propose a game-theoretic framework for understanding and improving multi-agent LLM systems in which an orchestrator decomposes tasks for a…
Researchers propose a game-theoretic framework for understanding and improving multi-agent LLM systems in which an orchestrator decomposes tasks for a…
A new study isolates how quantization interacts with recurrent neural network memory, introducing the concept of recurrent-state write-back to describe…
Researchers introduce Group Adaptive Clipping Policy Optimization (GAPO), a modification to Group Relative Policy Optimization methods used in reinforcement learning…
A new study investigates why reinforcement learning with verifiable rewards (RLVR) improves single-sample accuracy while narrowing a language model’s solution…
Researchers investigate whether the different functional operations that make up a language model’s chain-of-thought reasoning, such as problem formulation, goal…
Researchers introduce LLaDA-Image, an open framework pairing a 6B diffusion transformer with a frozen vision-language module built on the LLaDA2.0-Mini…
A new paper formulates conditional experience transfer for autonomous LLM post-training, addressing when past training-update evidence remains valid after a…
EarlyEval introduces early outcome prediction as a way to cut the cost of evaluating LLM agents, training lightweight classifiers to…
Researchers introduce Declarative Attention (DA), a protocol that has a language model explicitly declare in its chain-of-thought which parts of…
S3Gym is a new interactive benchmark testing whether LLM agents can turn accumulated experience into genuine self-improvement, evaluating three coupled…
A new study finds that real-world coding requests differ sharply from the curated GitHub issues used in SWE-bench-style benchmarks: 88%…
Cliff is a new reward-shaping method for reinforcement learning with verifiable rewards (RLVR) that uses an off-the-shelf LLM to identify…