Researchers including Wanshun Su, Yang Shi, and Feihu Liu, working across Northwestern Polytechnical University, Peking University, Alibaba Group, and Tsinghua University, introduce OmniPack, a training-free two-stage token compression framework for omni-modal large language models that process audio and video. The first stage applies modality-specific pre-LLM compression that removes structural redundancy through importance ranking and spatial coverage selection, merging discarded tokens into retained ones, while the second stage applies query-conditioned inner-LLM compression that refines tokens after multimodal interaction using textual relevance and cross-modal collaboration. Evaluated across five benchmarks and three model architectures, OmniPack preserves 95.6% of baseline accuracy while cutting computational operations to about 10% of baseline requirements.
