Puro-2B presents an open-source pretraining recipe for affordable language-model training, demonstrating a 1.5B-parameter model trained from scratch on up to 1.4 trillion tokens using consumer-grade RTX 5090 GPUs for under $6.9K, matching Qwen2.5-1.5B performance. The authors introduce a “Puro Cost Scaling Law” showing roughly $4.4K suffices to reach Qwen2-1.5B-level performance, and release the full training pipeline, data, code, and model weights.
