Intention Distillation (INDI) trains vision-language-action models using a frozen teacher vision-language model to distill behavioral intent into the action decoder, recovering multimodal intent representations at an intermediate decoder layer to organize action prediction alongside execution-progress signals rather than relying solely on behavior cloning. The method shows substantial improvements across simulation benchmarks and real-world robotic tasks.