LangChain optimized NVIDIA’s open Nemotron 3 Ultra model for agent tasks by tuning the harness — system prompts, tool descriptions, and middleware — rather than the model weights. Through iterative, trace-driven evaluation, they raised its Deep Agents suite score from 0.80 to 0.86, reaching near-parity with Opus 4.8 (0.87) at roughly one-tenth the cost per run ($4.48 versus $43.48). The writeup also notes the limits of harness tuning: once performance plateaus, further gains require post-training rather than more scaffolding changes.
