Netflix Details Its In-House LLM Serving Platform With Triton and vLLM
Netflix built an in-house LLM serving platform combining Triton and vLLM to handle inference across CPUs and GPUs within its…
Netflix built an in-house LLM serving platform combining Triton and vLLM to handle inference across CPUs and GPUs within its…
As collecting real-world robotics data remains expensive and slow, developers increasingly rely on GPU-accelerated simulators to generate synthetic training data…
An investigation detailed a marketplace, operating mainly out of China, for reselling discounted LLM API tokens using open-source proxy software…
Hugging Face reports successful testing of AMD’s new Instinct MI455X GPU, which offers 432 GB of high-bandwidth memory, more than…
The vLLM Semantic Router project outlines its evolution from routing requests between fast and reasoning model paths toward coordinating full…
vLLM describes the testing infrastructure it uses to maintain production quality, including a continuous integration system running 266 jobs across…
vLLM announced day-0 support for Thinking Machines Lab’s Inkling, a large multimodal mixture-of-experts model that accepts text, image, and audio…
LangChain explains how it evaluates its open-source Deep Agents framework using Harbor, an evaluation runner that requires specifying an agent,…
LangChain introduces an Eval Engineering Skill that helps coding agents automatically build evaluations by analyzing repository context and production traces.…
LangChain outlines a governance framework for enterprises deploying AI agents, arguing that LLM gateways should act as a runtime control…
Google Research explains how diffusion models generate novel images instead of memorizing training data. The research shows that this “creativity”…
In this post, researcher Lilian Weng examines how harness engineering, the systems that orchestrate model execution, tool use, and workflow…