Reward-Optimized Probe-and-Respond (RO-PnR) is a framework that decides when conversational agents should ask clarifying questions versus give direct corrections during health-misinformation interventions, using turn-level rewards to balance the value of gathering more information against interaction cost while modeling user variability through latent states representing health literacy and belief commitment. Experiments show RO-PnR achieves better cost-adjusted outcomes with 30% fewer conversational turns than always-probe strategies.