Agent-G2 addresses reward sparsity in long-horizon agentic RL by replacing the common practice of treating expert-trajectory guidance depth as a fixed scalar with a Gaussian distribution whose center and spread are estimated online from already-collected rollouts. This avoids the need for extra probe rollouts or a learned depth predictor while still capturing per-task heterogeneity in how much guidance is useful. On the ALFWorld and WebShop benchmarks with Qwen2.5 models, Agent-G2 outperforms the strongest hint-based, hint-free, and auxiliary-RL baselines by 2.3-7.4 points at under a third of the rollout cost of per-sample probing.
