StepGuard is a guard model that audits LLM-agent trajectories and monitors tool actions before execution, paired with StepGen, an automatic data-generation engine producing paired safe and unsafe trajectories, and Balance-GRPO, a training method that dynamically balances learning between safe and unsafe actions based on observed accuracy. Deployed on benchmark environments, StepGuard achieves significant security improvements while maintaining task utility.
