The study introduces a benchmark evaluating whether LLMs can construct and iteratively refine their own agent execution harnesses rather than just complete tasks directly. Agents first build a system from minimal specifications, then improve it using performance feedback, tested across six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 instances. Results show LLM-generated harnesses substantially underperform human-engineered ones in code and search domains while matching them in writing and ML experimentation, with iterative improvement gains proving inconsistent and largely failing to transfer to new models.
