HomeTechnologyOpenAI Unveils Ultrafast GPT-5.6 Sol at 14x Speed

OpenAI Unveils Ultrafast GPT-5.6 Sol at 14x Speed

New Cerebras-powered tier hits 750 tokens per second for real-time frontier AI.

OpenAI just removed a major barrier between top-tier intelligence and real-time use. On August 13, 2026, the company previewed Ultrafast mode for its flagship GPT-5.6 Sol model. The new API service tier runs the same model up to 14 times faster than standard processing and generates as many as 750 output tokens per second.

Powered by Cerebras wafer-scale hardware, Ultrafast keeps full Sol capability while delivering speeds previously limited to smaller or specialized models. OpenAI positions it as a way to get “more useful work per second” in time-critical settings.

What Changes in Practice

Standard GPT-5.6 Sol typically outputs around 50-plus tokens per second on conventional infrastructure. Ultrafast jumps that figure dramatically. Early internal tests at OpenAI show security investigations that once took one to two hours now finish in 10–15 minutes, sometimes approaching real time. Research loops that previously required overnight runs can now support multiple iterations in a single workday.

OpenAI lists clear use cases: live incident response, financial research and fraud checks, complex customer support and voice interactions, commerce personalization, and interactive coding or experimentation. The company is already testing the mode internally for log analysis, root-cause investigation, and rapid knowledge synthesis.

Early Customer Feedback

Select companies across coding, finance, commerce, and support have been running production workloads. Jane Street’s John Crepezzi noted the speed enables more focused developer workflows. Podium’s Courtland Lykins said it transforms the experience in voice AI stacks for complex calls. Basis and Rogo reported similar gains in synchronous user experiences and real-time financial research.

Cerebras benchmarks place Ultrafast roughly 11 times faster than Claude Fable 5 and five times faster than Claude Opus 4.8 in Fast mode, based on independent speed data. On Humanity’s Last Exam, the mode completed the full 2,500-question set in just over 11 hours versus more than three days for Fable 5 at comparable accuracy.

Availability and Pricing Context

Ultrafast launches first in the OpenAI API as a limited preview for a select group of customers. OpenAI will expand access as capacity grows and has opened a sign-up form for updates. No public pricing has been released yet. For comparison, standard Sol costs $5 per million input tokens and $30 per million output tokens. Fast mode, introduced earlier, delivers up to 2.5x speed at roughly double the standard price.

The move builds on OpenAI’s earlier partnership with Cerebras and the July general availability of the GPT-5.6 family (Sol, Terra, and Luna). It signals a shift: frontier models no longer force a hard trade-off between intelligence and latency.



RELATED ARTICLES

Most Popular

Recent Comments