Researchers propose MoTE, a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared, using sample-level task routing where each example follows a single task route rather than token-level routing. Evaluated on five COIN benchmarks, a five-expert VideoLLM-MoTE model activates about 2 billion parameters per sample and achieves higher average top-1 accuracy than recent VideoLLM baselines, outperforming both dense activation and learned sparse-routing alternatives.