OpenAI’s GPT-6 Astra Ultrafast mode is now available in the API and to eligible ChatGPT Work and Codex users, running on Nvidia Blackwell graphics processors. Nvidia says the mode generates tokens up to eight times faster than Astra Standard, with the largest practical benefit expected in repeated agent workflows.

A coding agent may alternate between writing code, calling a tool, reading the result and deciding what to do next. Reducing generation time at every turn can shorten that edit-test-debug cycle and make interactive applications respond more quickly. The claimed multiplier is a vendor figure and may vary with prompt length, output size, load and application design.

OpenAI and Nvidia also say models helped optimize the inference software that serves Astra on Blackwell hardware. Because the GPUs are programmable, the companies can test and deploy new kernels rather than treating inference performance as fixed after the model ships.

Developers can use Ultrafast through the OpenAI API now, with access and pricing details in OpenAI’s implementation guide. The announcement focuses on latency rather than changes to model intelligence or accuracy. Teams evaluating the mode should therefore measure complete task time and cost, including tool waits and network overhead, instead of assuming an eightfold improvement for an entire application.