Researchers built Daedalus-150M, a 150-million-parameter language model architected specifically for efficient CPU inference rather than adapted from a GPU-designed model. Its hybrid architecture uses full attention in 6 of 18 blocks and short convolutions in the remaining 12, capping convolution memory at just two timesteps regardless of conversation length. Trained on 59.9 billion tokens, it outperforms comparably sized models trained on far more data while decoding 1.76x faster at 2048-token context lengths and producing files 6.3% smaller than all-attention baselines.
