OpenAI is previewing a new API service tier called Ultrafast that is designed to run GPT-5.6 Sol at much higher speed. The company says the mode can generate up to 750 output tokens per second and operate up to 14 times faster than Standard processing.
The early version is powered by Cerebras and is launching first for the OpenAI API. OpenAI frames the tier as a way to bring frontier-model capability into time-sensitive products where users previously had to choose a smaller or more specialized model to get low latency.
The company points to coding, commerce, financial analysis, and other interactive business workflows as possible fits. Faster output can matter when a model is paired with a live user interface, an agent loop, or an application that must respond while a customer waits.
Access is still limited during the preview period. That matters because the post is less a broad product launch than a signal about model-serving strategy: labs are competing not only on intelligence and price, but also on how much useful work a strong model can deliver per second.