Researchers introduce BPCO, a reinforcement learning recipe combining decoupled PPO, bounded value predictions, Monte Carlo targets, and adaptive advantage estimation to stabilize critic-based training for language models. The method lets critics access reward-defining information such as reference answers during training while maintaining competitive performance. Experiments across models from 1.5B to 30B parameters show BPCO matches or exceeds group-based baselines while sampling only one response per prompt, positioning critics as a reliable alternative to group-relative advantage estimation.