OpenAI Previews Ultrafast Mode Running GPT-5.6 Sol at Up to 14 Times Standard Speed

OpenAI Previews Ultrafast Mode Running GPT-5.6 Sol at Up to 14 Times Standard Speed

A new Cerebras-powered API tier promises GPT-5.6 Sol responses at up to 750 tokens per second, though pricing remains undisclosed and access is limited to a select group of customers.

OpenAI announced on 13 August 2026 that it is previewing a new service tier called Ultrafast mode, which it claims can run its GPT-5.6 Sol model at up to 14 times the speed of its standard processing tier. That headline figure — 14 times faster — is the most aggressive performance claim OpenAI has publicly attached to any of its API tiers to date.

The announcement came via the official @OpenAI account on X, with the company stating: “Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.”

The mode is powered by Cerebras hardware — specifically, wafer-scale silicon chips that Cerebras has designed for large-scale AI inference. OpenAI and Cerebras have both confirmed the partnership. The maximum throughput figure cited is 750 output tokens per second, which OpenAI and Cerebras frame as a ceiling metric rather than a guaranteed rate across all workloads.

What Is GPT-5.6 Sol, and Why Does Speed Matter?

GPT-5.6 Sol sits within OpenAI’s GPT-5.6 model family, described in industry coverage as a successor to the earlier GPT-4 and GPT-5 series, with improved reasoning and coding performance. Ultrafast mode is not a new underlying model — it is a deployment configuration designed to reduce latency for applications where response time is a critical factor.

OpenAI positions the tier specifically at “products and workflows where every second counts.” That language points to use cases such as real-time customer support, high-frequency trading tools, live code analysis, and interactive simulations. These are contexts where waiting even a few seconds for a model response can break the user experience or, in financial applications, cost real money.

Cerebras’s wafer-scale engine approach is central to how those speeds are achieved. Unlike conventional AI accelerators built from multiple chips connected together, wafer-scale designs integrate processing across a single large silicon wafer, reducing the communication overhead that can slow inference down. The Ultrafast announcement formalises what had been mentioned in passing during earlier GPT-5.6 rollout communications.

The Numbers, and the Caveats

The figures are striking. But independent analysts are urging caution. The “up to 14x” and “up to 750 tokens per second” claims are vendor-stated maximums with no publicly published baseline workload, prompt mix, or quality trade-off data attached to them. There is no independently audited benchmark to validate either figure.

One developer-focused blog reported that GPT-5.6 Sol Ultrafast completed the Humanity’s Last Exam benchmark — a demanding, multi-subject evaluation used in AI research communities — in 11 hours and 11 minutes, compared with 78 hours and 27 minutes for Anthropic’s Claude Fable 5 on the same test. That comparison has not been verified by official data from either OpenAI or Anthropic, and should be treated accordingly.

The advice from technical commentators is consistent: treat the headline numbers as marketing ceilings, and test performance against your own prompts and workloads before drawing conclusions.

Pricing is another open question. As of mid-August 2026, OpenAI has not published a rate card or contract terms for Ultrafast mode. That absence matters. Faster inference typically requires more infrastructure, and there is a reasonable expectation that Ultrafast will cost more than the standard tier — possibly considerably more. Without published pricing, smaller businesses and public-sector bodies cannot assess whether the tier is financially viable for them.

Access Is Limited — and the Gap May Widen

Right now, Ultrafast is available only to a select group of OpenAI API customers. OpenAI says access will expand “to more businesses as capacity grows,” but no timetable has been given. It’s an early preview, and capacity constraints are driving the pace.

Critics point out that select-preview access tends to favour larger or strategically important enterprise customers, at least initially. That could widen the gap between well-funded organisations with existing OpenAI relationships and smaller firms that lack the leverage to get early access. For developers and start-ups hoping to build competitive real-time products, that gap is not trivial.

There are also broader safety questions. Faster models that generate outputs in near-real time raise the question of whether human oversight can keep pace. Some commentators have noted that speed at scale could amplify risks — rapid generation of disinformation, for instance, or automated code exploits — if safeguards are not matched to the throughput. OpenAI has not published specific safety documentation for the Ultrafast tier as of the announcement date.

The UK Government has not issued any statement specific to Ultrafast mode. However, existing frameworks from the Information Commissioner’s Office and the wider UK AI strategy apply to any deployment of the technology, including high-speed tiers. Organisations using Ultrafast for data-processing tasks would need to comply with UK data protection law as they would with any other AI service.

What This Means for Kent Residents

For most Kent businesses and residents, Ultrafast mode will not be directly accessible in the near term — the preview is limited to a select group of API customers, and no general availability date has been set. But if local firms in sectors such as logistics, retail, or customer services are already using the OpenAI API, they may find it worth registering interest with OpenAI as capacity expands. The bigger practical question — which cannot yet be answered — is what Ultrafast will cost, since undisclosed pricing makes it impossible for smaller Kent organisations, or public bodies such as NHS Kent and Medway Integrated Care Board, to plan meaningfully around the tier. UK consumers may eventually feel the effects indirectly, through faster AI-powered tools in apps and services they already use, if providers integrate Ultrafast into their products.

Source: @OpenAI

OpenAI Previews Ultrafast Mode Running GPT-5.6 Sol at Up to 14 Times Standard Speed Quiz

5 questions