The paper details training Nemotron 3 Ultra with supervised fine-tuning and reinforcement learning to produce natural-language mathematical proofs, built around an iterative test-time inference pipeline with three specialized checkpoints that generate, verify, and refine candidate proofs without relying on formal verification tools or external resources. The resulting system scored 30 of 42 points at IMO 2026, meeting the gold-medal threshold. The team is releasing the trained checkpoints, training code, and a new 200-problem olympiad-level benchmark, giving other practitioners a reproducible recipe for this class of reasoning fine-tuning.