Researchers present X-AuT, a technique for making speech-focused large language models more efficient by compressing their audio encoders through progressive layer selection and representation alignment. The method combines cross-scale distillation, scheduled student-policy supervision and LoRA finetuning while keeping the language model backbone frozen. Applied to Qwen3-ASR-0.6B, reducing audio-encoder layers from 18 to 16 improved performance across ten Chinese-English benchmarks, and further compression to 14 layers achieved competitive results with 20.7% fewer parameters in the audio tower.