Day 0 Support for Qwen3.8-2.4T-A95B on vLLM
vLLM announced Day 0 support for Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse mixture-of-experts model with 512 experts, running without architecture modifications. The…
vLLM announced Day 0 support for Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse mixture-of-experts model with 512 experts, running without architecture modifications. The…
vLLM reached 25,000 tokens per second per GPU serving Qwen3.5 on Blackwell GPUs via a disaggregated prefill-decode architecture. Three optimizations…
txtai adds LEMUR and mean-centering techniques for late-interaction retrieval. LEMUR learns a fixed-dimensional encoding of multi-vector representations so late-interaction models…
VIDRAFT’s AX-Ray diagnostic framework evaluates AI models for deployment safety beyond standard capability benchmarks, using 117 structured diagnostic items across…
Researchers discovered a security vulnerability in proprietary LLM APIs that allowed encrypted chain-of-thought reasoning traces to be extracted and read.…
Simon Willison describes Doug Turnbull’s technique for tagging blog content without constraining a model to a fixed vocabulary. Rather than…
Researchers from UIUC, the University of Maryland, Nanyang Technological University, Purdue, and the University of Illinois Chicago introduce LLMRouter, a…
DarwinX proposes evolving the harness around a frozen LLM — its prompts, tools, memory, and control flow — using a…
AutoDesign targets systems that turn multimodal content into structured design outputs, arguing that most such pipelines stay static rather than…
This paper proposes the Spatial Memory Agent (SMA), a framework that improves spatial reasoning in frozen vision-language models without any…
This research examines internal activation patterns in hybrid linear-attention language models and identifies two distinct phenomena: sharp pre-attention spikes that…
The paper identifies a bottleneck in autoregressive transformers where only sampled output tokens are fed back into the model, discarding…