CAST is a framework that addresses reliability challenges in LLM agents operating in long-horizon environments by converting sparse task outcomes into action-level supervision for critique learning and policy optimization. The method synthesizes structured rationales that explain action validity, then uses a trained critique model to generate supervision signals for policy fine-tuning. Applied to Qwen3 models on dynamic tool-calling benchmarks, CAST surpassed GPT-OSS-120B by more than 10 percent on retail tasks and achieved a further 9 percent gain on out-of-domain telehealth evaluations.