Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
WeiboAI researchers introduced CLR (Claim-Level Reliability Assessment), a training-free framework that improves LLM reasoning accuracy by directing verification effort at…
WeiboAI researchers introduced CLR (Claim-Level Reliability Assessment), a training-free framework that improves LLM reasoning accuracy by directing verification effort at…
Researchers released PRM-as-a-Judge 1.5, an updated toolkit that evaluates robotic manipulation policies by converting video rollouts into fine-grained progress metrics…
Researchers at the University of Milan-Bicocca built ParliamentRAG, a retrieval-augmented generation system for Italian Chamber of Deputies records that scores…
A new benchmark called MobileMem evaluates AI agents’ ability to provide persistent personal assistance by synthesizing a full year of…
Researchers introduced Mobius-v0, an architecture that separates knowledge storage from reasoning by using a globally shared memory implemented as a…
Researchers released Mimir v1, a 1-billion-parameter language model built on the Hierarchical Reasoning Model architecture and trained exclusively on ethically-sourced,…
A study examines how many times high-quality domain-specific data should be repeated during LLM pretraining as models and token budgets…
Dion3 targets the computational bottleneck of orthogonalizing weight updates in the Muon optimizer, introducing a Gram Newton-Schulz algorithm, specialized kernel…
Nathan Lambert’s Interconnects.ai analyzed how Z.ai’s GLM-5.3 model reaches frontier-level agentic coding performance without changing its roughly 750-billion-parameter base model…
The vLLM project detailed Distributed Layerwise Offload, a technique that lets large diffusion transformer models run across multiple GPUs or…
Simon Willison tested Alibaba’s Qwen 3.8 27B, a 27-billion-parameter vision-capable open-weight model that fits in a 17GB file, and found…
vLLM implements adaptive verification in speculative decoding using DSpark’s confidence-scored drafting mechanism, which dynamically adjusts verification length per step instead…