A paper proposes UniME-R1, a multimodal retrieval framework that generates reasoning conditioned on retrieval feedback rather than the query alone. Its embedder-adviser architecture analyzes initially retrieved candidates to spot discriminative cues the embedding model missed, then either reranks results or produces a retrieval-centric chain-of-thought to refine the search. The framework is trained by mining hard negatives to simulate realistic retrieval failures, combining supervised learning with retrieval-oriented reinforcement learning, and the authors report consistent gains across multiple multimodal retrieval benchmarks.
