Efficient Decode Context Parallelism with vLLM for Long Context Workloads
vLLM has introduced Decode Context Parallelism (DCP), a technique that splits key-value caches across GPUs by sequence dimension rather than…
vLLM has introduced Decode Context Parallelism (DCP), a technique that splits key-value caches across GPUs by sequence dimension rather than…
TNG Technology Consulting extended NVIDIA’s Nemotron 3.5 Lightning model with vision capabilities using a simplified approach designed for limited computing…
Meta and the PyTorch team detailed how ExecuTorch now runs Muse Glimmer, an open-weight 30-billion-parameter model distilled for on-device agentic…
Google DeepMind’s WeatherNext Cyclones model uses Functional Generative Networks to efficiently generate large probabilistic ensembles for tropical cyclone track and…
OpenAI found that its GPT-5.6 Sol model scored only 13.3% on the ARC-AGI-3 benchmark’s public task set when run through…
The Navid AI team investigated suspected vote manipulation on its Arabic TTS Arena leaderboard after one model won 91 of…
Hugging Face and EleutherAI created the FineBooks BHL OCR Leaderboard to test whether open-source OCR models can accurately digitize historical…
OpenAI describes a new feature letting Codex Code Review read custom repository rules from AGENTS.md files so it can catch…
Simon Willison walks through a published technical timeline of how OpenAI’s autonomous training-run agents, beginning May 7, 2026, escalated from…
Meta engineers describe a multi-stage sequence-modeling architecture for ads ranking that decouples heavy offline user modeling from lightweight, latency-sensitive online…
Meta details the hardware-software co-design behind doubling training efficiency for GEM, its LLM-scale ads recommendation foundation model, including custom recommendation…
After a Linux kernel upgrade caused latency regressions in its ads-serving infrastructure, Meta deployed sched_ext, an open-source BPF-based kernel scheduler,…