Researchers introduce tau-squared-Bench, a benchmark that evaluates whether AI coding agents can build complete, production-ready customer-service agents under realistic client-engagement conditions, rather than testing narrow coding tasks in isolation. Across 53 tasks spanning four business domains, the strongest tested configuration, Claude Opus 5 running under Claude Code, passed only 23.9% of evaluation simulations, compared to an 82.2% ceiling set by an expert-authored reference solution. The study found that failing agents issued shallow database queries instead of deeply understanding business records, communicated poorly with simulated clients, and shipped their first working design without experimenting with architecture or cost trade-offs.