This paper runs a controlled comparison of next-chunk reasoning reinforcement learning against a simple alternative called Mixed SFT — a single supervised fine-tuning stage that jointly trains on chain-of-thought and non-chain-of-thought data — for leveraging text corpora that lack explicit reasoning annotations. Mixed SFT achieves a clearly higher post-RL performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute, with the advantage holding across both in-domain math reasoning and out-of-domain tasks. The authors also find that higher accuracy before reinforcement learning does not reliably predict higher accuracy after it, arguing that no-CoT training strategies must be evaluated within the full post-training pipeline.