Cliff is a new reward-shaping method for reinforcement learning with verifiable rewards (RLVR) that uses an off-the-shelf LLM to identify the first mistake in a reasoning rollout, then assigns positive advantage to the correct prefix and negative advantage to everything after it. Across 12 evaluation scenarios, Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even when the teacher model identifying mistakes has only modest capability.