OpenAI introduced a preview of “Ultrafast,” a new API service tier that runs its GPT-5.6 Sol model on Cerebras hardware at up to 750 output tokens per second, roughly 14 times faster than the standard processing tier. The speed boost draws on a multiyear partnership announced in January under which OpenAI plans to deploy up to 750 megawatts of Cerebras inference systems through 2028, reportedly worth more than $10 billion. Ultrafast preserves the same underlying intelligence as the standard GPT-5.6 Sol model and is aimed at time-sensitive use cases such as incident response, financial analysis, and customer support. The tier is currently limited to select API customers, with pricing and general availability still to be announced.
