OpenAI described engineering work, partly performed autonomously by its GPT-5.6 “Sol” model, to optimize the production software that serves its models, reducing end-to-end serving costs by about 20%. The effort also improved speculative decoding techniques, increasing token-generation efficiency by more than 15%, alongside better hardware-utilization routing and context-management changes designed to prevent agents from repeating work. Within a human-supervised process, Sol rewrote and optimized production kernels and designed and ran hundreds of experiments to improve token generation.