Researchers present CalibForge, a system that automatically synthesizes and calibrates training tasks for terminal-use AI agents based on verified solver behavior. It applies two calibration strategies, one targeting disagreement across a pool of diverse solvers and another targeting a designated strong-pass/weak-fail relation, to keep each task within a solver-relative learnable difficulty zone. Using the method, the team built more than 5,400 calibrated tasks and reports gains of up to 24.71 percentage points on Terminal-Bench 2.0 compared with tasks produced by authoring and validation alone.