Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense
This guide details how to deploy the open-weight GLM 5.2 model on enterprise infrastructure for cybersecurity incident response, arguing that…
This guide details how to deploy the open-weight GLM 5.2 model on enterprise infrastructure for cybersecurity incident response, arguing that…
POCKET is a 34.66-billion-parameter sparse Mixture-of-Experts model engineered for efficient on-device inference by activating only about 3 billion parameters per…
PyTorch’s Helion, a domain-specific language for performance-portable ML kernels, now compiles to optimized TPU code via Pallas using a two-level…
This technical report examines automated LLM evaluation systems, arguing the field has shifted from reference-based metrics like BLEU and ROUGE…
The vLLM AFD plugin introduces Attention-FFN Disaggregation, separating a Mixture-of-Experts model’s attention and FFN components into independently deployed services that…
The vLLM team details engineering work to serve Moonshot AI’s Kimi K3 model at production scale, including a prefix-caching design…
Google announced three new AI models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash…
Block, the payments company led by Jack Dorsey, launched Buzz on July 22, 2026, a free open-source platform built on…
SkyPilot, a startup cofounded by Databricks cofounder Ion Stoica and Zongheng Yang, announced on July 21, 2026 that it raised…
Chinese robotics firm UBTECH unveiled its U1 Series on July 1, 2026, a full-sized consumer humanoid robot positioned as an…
Microsoft and Mistral announced an expanded strategic partnership on July 21, 2026 aimed at bringing frontier AI to enterprises and…
Amazon announced job cuts within its Artificial General Intelligence organization on July 22, 2026, as part of an effort to…