This paper presents a production-driven methodology for consolidating fragmented enterprise LLM deployments, including a template-aware sampling technique for building internal benchmarks from production traffic and a modular post-training recipe that trains separate GRPO experts for instruction-following, function-calling, and dialogue before merging them via two-stage SLERP. The authors identify and fix three distinct reward-hacking failure modes along the way. The resulting 32B model matches a seven-times-larger baseline while serving 116 million monthly requests.