LangChain introduces an Eval Engineering Skill that helps coding agents automatically build evaluations by analyzing repository context and production traces. The skill proposes testable properties through an iterative interview process with the user before generating executable evaluations in Harbor format. The author argues that effective eval design needs human feedback rather than one-shot generation, and that containerized, reproducible test environments let teams mine production data for recurring failures and continuously improve their agents.
