vLLM announced day-0 support for Thinking Machines Lab’s Inkling, a large multimodal mixture-of-experts model that accepts text, image, and audio input with up to 1 million tokens of context. The implementation reports throughput of 380 tokens per second per user with multi-token prediction, using optimizations such as sconv-aware tensor parallelism and low-latency fused collective operations. vLLM says accuracy checks across vision, audio, tool-calling, and long-context benchmarks confirm its implementation matches the reference model’s performance.
