OpenAI is making inference speed a distinct part of its API product with a new Ultrafast mode for GPT-5.6 Sol, The Decoder reports.

The mode is powered by Cerebras hardware and is described as delivering up to 750 output tokens per second. It sits alongside Standard and Fast options, creating a three-tier structure where customers can pay for different balances of speed and cost.

The practical change is that latency becomes easier to buy explicitly. For interactive coding tools, voice systems, agents, and customer-facing apps, response time can matter as much as raw model quality. A faster tier lets developers reserve the highest throughput for workflows where delays are most visible, while using slower options for background jobs.

The report also shows how model services are becoming infrastructure products with several dimensions: capability, price, reliability, and now speed. The limitation is that advertised token rates do not automatically translate into better end-user experiences. App design, network latency, tool calls, and prompt length still shape how fast a system feels.