This study measures the cost-quality trade-off of switching between cheaper and more capable models partway through a long-running coding agent task, using paired low-cost and high-cost models from the Claude and GPT families. Escalating from a cheaper to a stronger model recovers less than half of the quality gap between them while incurring a substantial cost premium — a penalty the authors term the ‘handoff tax’ — whereas downshifting from a strong to a cheap model offers a more favorable cost-quality trade. The paper also finds the ideal amount of trajectory information to hand off reverses depending on direction: trimming the weaker model’s trajectory helps escalation, while trimming the stronger model’s trajectory hurts downshifting.
