AgentJudgeBench is a benchmark of 3,808 synthetically generated records spanning six DAG topologies and three difficulty tiers, built to systematically test the reliability of LLM judges on agentic tool-calling tasks. It uses a paired with/without-ground-truth evaluation protocol checked against a deterministic programmatic scorer, spanning 321,648 paired evaluations across five generator and six judge models. The authors report findings on how difficulty, ground-truth exposure, temperature, chain-of-thought, and prompt format affect judge alignment.
