OpenAI has previewed a new service tier called Ultrafast that is designed to run GPT-5.6 Sol at substantially lower latency than its standard processing option. The company says the new tier can operate at up to 14 times the speed of Standard processing and generate as many as 750 output tokens per second.

Ultrafast is launching first through the OpenAI API and is currently available only to a limited group of customers. OpenAI says it plans to broaden access as capacity grows.

The service is powered by Cerebras, extending an existing partnership between the two companies around low-latency inference. Rather than pairing faster responses with a smaller model, OpenAI is using the infrastructure to serve GPT-5.6 Sol itself at higher speed.

That distinction is central to the product. Faster inference has often required developers to choose a smaller or more specialized model when response time is critical. Ultrafast is intended to reduce that tradeoff by making a frontier model usable in workflows where seconds — or fractions of seconds — can materially affect the experience.

OpenAI highlights incident response, live research, customer support, voice applications, commerce and financial analysis as examples of workloads that could benefit. In those settings, the value of higher throughput is not simply a faster chat response. It can shorten an entire loop of reading information, reasoning over it, producing an answer and moving to the next action while the underlying situation is still changing.

The launch also turns inference speed into a more visible product dimension for frontier AI systems. Model comparisons have traditionally focused heavily on capability, benchmark performance and price. As advanced models become useful enough for interactive operational work, latency and sustained generation speed are becoming increasingly important parts of the deployment decision.

For now, Ultrafast remains a preview rather than a generally available API tier. Its broader impact will depend on how widely OpenAI can expand capacity and whether developers see enough practical benefit from the additional speed to redesign latency-sensitive workflows around it.