Verification-Aware Training Improves Speculative Decoding Speed for LLMs
Researchers introduce Verification-Aware Training (VAT), a framework that improves speculative decoding for large language models. VAT adds a lightweight verification…
Researchers introduce Verification-Aware Training (VAT), a framework that improves speculative decoding for large language models. VAT adds a lightweight verification…
CAST is a framework that addresses reliability challenges in LLM agents operating in long-horizon environments by converting sparse task outcomes…
Researchers introduce StudentSim, a training framework that creates individualized student simulators by combining pooled training with per-student specialization. The accompanying…
Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that integrates 3D perception, visual question answering, and motion planning into…
SMELT loops the middle half of layers twice in Mixture-of-Experts Transformers while holding FLOPs, parameter count, and KV cache constant.…
UI-Venus-2 is a multimodal GUI agent built to automate digital tasks across mobile, web, and desktop environments through unified reasoning…
H3-World adapts the MiniMax-H3 video generator into an interactive world model by converting character and camera actions into compositional natural-language…
ZimaBlue trains World Action Models from large-scale egocentric video through a three-stage curriculum: causal video pre-training, grounding in robot trajectories,…
This paper presents a production-driven methodology for consolidating fragmented enterprise LLM deployments, including a template-aware sampling technique for building internal…
Hi-Q is an evidence-conditioned framework for multi-hop question answering that dynamically refines queries into hierarchical trees by testing whether retrieved…
This paper introduces DroneCATS-Agent, an architecture letting multimodal LLMs control drones directly through natural-language prompts, and DroneCATS, a benchmark spanning…
This paper studies how visual understanding and generation tasks interact inside unified multimodal models through controlled experiments. The authors propose…