ZimaBlue trains World Action Models from large-scale egocentric video through a three-stage curriculum: causal video pre-training, grounding in robot trajectories, and target-robot specialization. A Slow-Fast dual-system architecture balances model capacity against real-time control demands. Incorporating over 120,000 hours of embodied video lifts zero-shot task success on real-robot evaluations from 36.1% to 77.8%.