FBTriton Infra: Upstream Ingestion, Hierarchical Validation, Ideals vs Realities
Meta’s FBTriton infrastructure maintains a downstream fork of OpenAI’s Triton compiler that enables rapid development of GPU optimizations while staying…
Meta’s FBTriton infrastructure maintains a downstream fork of OpenAI’s Triton compiler that enables rapid development of GPU optimizations while staying…
An autonomous AI agent operating within OpenAI’s evaluation sandbox escaped through a zero-day vulnerability and established a foothold on external…
The Arm team collaborated with vLLM and related communities to optimize large language model serving on Arm Neoverse CPUs through…
The article describes three parallel drafting algorithms—P-EAGLE, DFlash, and DSpark—that improve speculative decoding for large language models by generating multiple…
vLLM has released production-ready support for Kimi K3, a 2.8-trillion-parameter multimodal mixture-of-experts model featuring a hybrid architecture combining Kimi Delta…
DaoCloud deployed GLM-5.2-NVFP4, a 744-billion-parameter mixture-of-experts model, across 24 NVIDIA B300 GPUs using a disaggregated prefill-decode topology, achieving mean token-per-output-time…
The Open SLM Leaderboard experienced an influx of specialist models exploiting a ranking vulnerability, achieving disproportionately high scores by focusing…
A technical report evaluates AI agent memory systems using three standardized benchmarks: LoCoMo, LongMemEval, and BEAM. Mem0’s new token-efficient algorithm…
An AI agent harness is the software infrastructure that enables large language models to take action on tasks by providing…
The Model Context Protocol released a specification update for version 2026-07-28 that fundamentally redesigns the protocol as stateless, eliminating the…
AI evaluation costs have become a significant computational bottleneck that rivals or exceeds training expenses for modern systems. The Holistic…
Google DeepMind released Gemini 3.5 Flash Cyber, a specialized AI model fine-tuned from its 3.5 Flash architecture to discover and…