Published July 30, 2026, this article describes how researchers achieved significant inference speedups running the Qwen3.6-35B-A3B model on Intel Core Ultra Series 3 processors using DFlash speculative decoding, which proposes tokens in parallel blocks rather than sequentially. The technique delivered a 2.2x speedup on coding tasks, 1.6x on math, and 1.3x on chat benchmarks for the MoE model, with speedups reaching 5.5x on dense Qwen3.6-27B variants. The results demonstrate practical acceleration across both mixture-of-experts and dense Qwen models for local AI deployment on consumer hardware.
