OpenAI previews Ultrafast mode: GPT-5.6 Sol up to 14x faster

OpenAI has shared an early look at Ultrafast, a new service tier that runs its GPT-5.6 Sol model up to 14 times faster than standard processing. The tier is powered by Cerebras and generates up to 750 output tokens per second. It is launching first inside the OpenAI API, in a limited preview available today to a select group of customers, with access expanding as capacity grows.
OpenAI frames Ultrafast as a new speed class rather than a smaller or specialized model: previously, getting real-time responsiveness meant trading down to a less capable model. Ultrafast is meant to deliver Sol's full intelligence at that speed instead. The company lists five scenarios it has seen work well during early testing: incident response and reliability (reading logs, code changes and engineer reports to find a likely cause while an outage is still unfolding), financial research and security (analyzing market signals and transactions for suspicious activity as conditions change), customer support and voice (resolving multi-step issues without breaking the flow of a live conversation), commerce (answering product questions and resolving checkout issues while a shopper is still deciding), and live research and experimentation (turning overnight batch runs into interactive working sessions).
OpenAI says it has been running its own internal group of developers on GPT-5.6 Sol with Ultrafast mode to see which workflows benefit most. One example given is incident response: engineers use it to read logs, analyze traces, synthesize conversations and help prepare or validate a fix while the evidence is still changing, with OpenAI noting that engineers remain responsible for judgment and deployment. A second internal example is research, where the model is used to search knowledge sources and summarize information across tools; OpenAI says this is tightening the usual overnight-batch-then-morning-review loop into something that can support multiple iterations within a single workday.
Ultrafast is described as the next step in OpenAI's partnership with Cerebras on ultra-low-latency inference. The post does not disclose pricing for the tier, a specific date for when the preview began or when it will reach general availability beyond "today" and "as capacity grows," an input/prompt processing speed figure (only the output rate is given), or the names of the customers in the initial testing group. Businesses that want frontier-level intelligence at this speed can sign up to be notified as access expands.
Key facts
- Ultrafast runs GPT-5.6 Sol up to 14x faster than OpenAI's standard processing, generating up to 750 output tokens per second.
- The tier is powered by Cerebras hardware and launches first in the OpenAI API, in limited preview to a select group of customers today.
- OpenAI points to five use cases from early testing: incident response, financial research and security, customer support, commerce, and live research.
- OpenAI's own internal teams have been using Ultrafast for incident response and research, tightening an overnight batch-and-review loop into same-day iteration.
- No pricing, no general-availability date, no input-processing speed figure, and no customer names have been disclosed; interested businesses can sign up to be notified as access expands.
Why it matters
Ultrafast is OpenAI's attempt to stop speed and intelligence from being a trade-off. Until now, getting real-time responsiveness from an OpenAI model typically meant dropping to a smaller or more specialized model. Ultrafast instead runs the full GPT-5.6 Sol model up to 14x faster, at up to 750 output tokens per second, which OpenAI positions as opening up time-sensitive work that previously required compromising on model capability.
Who it affects
The preview is aimed at businesses with time-sensitive workflows: incident response and reliability teams, financial research and security teams, customer support and voice products, commerce and checkout flows, and research teams running iterative experiments. Access is currently limited to a select group of customers OpenAI is working with directly, plus OpenAI's own internal developers, who have been testing the mode on incident response and research workflows.
How to use it
Ultrafast is launching first in the OpenAI API and is in limited preview today, available only to a select group of customers as OpenAI studies real production use. There is no public pricing yet and no announced date for wider availability beyond OpenAI's statement that access will expand "as capacity grows." Businesses that want frontier intelligence at this speed can sign up to be notified when access opens further.
How solid is it
The claims come directly from OpenAI's own announcement, with two concrete, specific figures: up to 14x the speed of standard processing and up to 750 output tokens per second, attributed to hardware from Cerebras. OpenAI also describes internal use by its own engineering and research teams as supporting evidence, though this remains a company blog post about the company's own product rather than an independently verified benchmark.
Risks and caveats
This is a limited, invitation-based preview, not a general release, so most businesses cannot access it yet. OpenAI has not disclosed pricing for the tier, a firm timeline for broader availability, an input or prompt-processing speed figure (only the output generation rate is given), or the names of any customers in the initial testing group. The performance figures and use-case examples come from OpenAI itself rather than from an independent source.