Tencent Hunyuan researchers introduce CAFE, a framework in which a shared-parameter model alternates between agent and critic roles to build self-improving search systems, combining online reinforcement learning using comparative feedback estimates with offline preference optimization from matched trajectories. Evaluated across seven agentic search benchmarks, CAFE outperforms other RL-based agents, maintains gains on six out-of-domain tests, and reduces hallucinations. The authors show that alternating the two updates continues improving performance compared to optimizing either component alone.
