OpenAI previews Ultrafast tier for GPT-5.6 Sol, up to 14X faster

OpenAI previews Ultrafast tier for GPT-5.6 Sol, up to 14X faster

OpenAI has announced an early preview of Ultrafast, a new service tier that runs its GPT-5.6 Sol model up to 14 times faster than Standard processing. Powered by Cerebras, Ultrafast generates up to 750 output tokens per second. The tier is launching first in the OpenAI API and is available today only to a limited group of preview customers; OpenAI says it will expand access as capacity grows, and interested businesses can sign up to be notified.

OpenAI frames Ultrafast as a new speed class rather than a smaller or cheaper model: until now, getting real time speed typically meant choosing a smaller or more specialized model instead of the most capable one. Ultrafast is meant to let businesses keep the full intelligence of GPT-5.6 Sol while operating at real time speed. OpenAI lists five scenarios where it expects Ultrafast to matter: incident response and reliability, where a system analyzes logs, recent code changes, and engineer reports to help find a cause and prepare a fix while an outage is still unfolding; financial research and security, analyzing market signals and transactions to flag suspicious activity while conditions are still changing; customer support and voice, resolving complex, multi-step customer issues in real time without breaking the conversation; commerce, answering product questions and resolving checkout issues while a shopper is still deciding; and live research and experimentation, turning what used to be an overnight batch run into an interactive session where a team can test an idea, check the results, and run another experiment without losing momentum.

OpenAI says it has already been testing GPT-5.6 Sol on Ultrafast mode internally, alongside an initial group of external companies across coding, commerce, financial research, and support. Inside OpenAI, one internal use case is incident response: when an alert fires, teams use Ultrafast to quickly read logs, analyze traces, synthesize conversations, identify the next checks, and help prepare or validate a fix, all faster, while engineers remain responsible for judgment and deployment. A second internal use case is research: team members previously launched a batch of experiments overnight and reviewed results the next morning; with Ultrafast, OpenAI says that loop is tightening to support multiple iterations during the workday instead of one review the next day.

OpenAI describes Ultrafast as the next step in its partnership with Cerebras on ultra low latency inference. The announcement gives no pricing or cost details, no named individual is quoted or credited, no specific preview customers are named, and no timeline is given for when Ultrafast will move from limited preview to general availability.

Key facts

  • Ultrafast runs GPT-5.6 Sol up to 14X faster than Standard processing and generates up to 750 output tokens per second.
  • The tier is powered by Cerebras hardware, marking the next step in OpenAI's partnership with Cerebras on ultra low latency inference.
  • Ultrafast is launching first in the OpenAI API and is available today only in a limited preview to a select group of customers; OpenAI says access will expand as capacity grows.
  • OpenAI highlights five target use cases: incident response and reliability, financial research and security, customer support and voice, commerce, and live research and experimentation.
  • OpenAI's own teams already use Ultrafast internally for incident response and research, saying the research loop is tightening from an overnight batch reviewed the next morning to multiple iterations within a single workday.

Why it matters

Until Ultrafast, OpenAI says getting real time speed typically meant choosing a smaller or more specialized model instead of the most capable one. Ultrafast is positioned as a new speed class that keeps the full intelligence of GPT-5.6 Sol while running up to 14X faster than Standard processing, at up to 750 output tokens per second, powered by Cerebras. That combination is what OpenAI is calling a competitive advantage: intelligence and speed no longer traded off against each other.

Who it affects

OpenAI points to businesses with time sensitive workflows: teams doing incident response and reliability work, financial research and security analysis, customer support and voice interactions, commerce and checkout flows, and live research or experimentation. For now, access is limited to an initial group of external preview customers spanning coding, commerce, financial research, and support, plus OpenAI's own internal teams, who are already using Ultrafast for incident response and research.

How to use it

Ultrafast is launching first in the OpenAI API and is available today only in a limited preview to a select group of customers. OpenAI says it will expand access as capacity grows, and businesses that want frontier intelligence at the highest speed can sign up to be notified when access expands. The announcement gives no pricing or cost information for the tier.

How solid is it

The information comes directly from OpenAI's own announcement, with two concrete figures: up to 14X the speed of Standard processing and up to 750 output tokens per second. Both are stated as maximums ('up to'), not guaranteed or average throughput. No independent benchmark is cited, no named executive or engineer is quoted, and no specific external preview customers are identified.

Risks and caveats

Because the headline figures are best case maximums, actual throughput for a given workload may be lower. No pricing is disclosed, so the cost of Ultrafast relative to Standard processing is unknown. Access is currently limited to a select group of customers with no stated timeline for general availability. For the incident response use case, OpenAI notes that engineers remain responsible for judgment and deployment even when Ultrafast speeds up the investigation.