Researchers introduced FACET, a framework for generating training tasks for terminal/command-line agents that keeps instructions, solutions, and verifiers aligned by grounding all task artifacts in a shared, executable environment state. The system reconstructs agent skills into information-rich scenarios and combines execution-based validation with targeted repairs to produce complex tasks with dense verification checks. FACET showed improved performance across multiple model scales on the Terminal-Bench 2.1 benchmark.
