This paper presents a cross-lingual evaluation of extractive prompt compressors across ten languages and eleven target models, benchmarking four learned compressors and four deterministic baselines on fully parallel data with tokenizer-matched budgets across more than 250,000 evaluation calls. The study finds English-trained compressors degrade significantly in non-English languages, while multilingually trained compressors perform comparably across languages, revealing whether compression narrows or widens the token-cost gap between English and other languages.
