Researchers introduced MMDiff, a framework for discovering and controlling interpretable features in multimodal large language models by comparing sparse autoencoders trained on a base language model against its vision-language-adapted counterpart. The method isolates features altered by multimodal training through geometric rotation analysis and visual-responsiveness metrics, then narrows to task-specific feature subsets using contrastive token firing with lexical-invariance filtering. Across three MLLM families and tasks spanning spatial reasoning, safety, and OCR, ablating discovered features degraded targeted capabilities by 12-17% while preserving general visual question-answering performance, and steering with the identified features improved spatial and OCR accuracy by 3.6% and 1.8% respectively over baselines.
