OpenAI's 'Ultrafast' Mode Hits 750 Tokens/Sec, Powered by Cerebras Wafer-Scale Chips
GPT-5.6 Sol now runs 14x faster without sacrificing intelligence. Here's why unbundling speed from model size changes the math for agentic workflows.
For the entire modern AI era, developers have been forced into a frustrating compromise: if you want frontier-level intelligence, you have to wait for it. If you need real-time speed, you have to downgrade to a smaller, distilled, or specialized model.
As of this week, that trade-off is dead.
On August 13, OpenAI introduced Ultrafast, a new API mode that runs its flagship GPT-5.6 Sol model at a blistering 750 output tokens per second. That is roughly 14 times the speed of standard processing, and it fundamentally alters the economics and user experience of building with LLMs.
Crucially, OpenAI hasn't achieved this through quantization or model distillation. The intelligence remains completely intact. This is the exact same frontier GPT-5.6 Sol model—just running on entirely different silicon.
Here is a deep dive into how OpenAI shattered the memory-bandwidth wall, why they partnered with Cerebras to do it, and why this marks the moment that "speed" officially becomes a productized tier in the AI wars.
The Specs: What 750 Tokens Per Second Actually Means
To put 750 tokens per second into perspective, the average human reads at about 4 to 5 tokens per second. At 750 t/s, GPT-5.6 Sol is generating text roughly 150 times faster than you can read it. It can output a standard 2,000-word essay in less than four seconds.
According to OpenAI's preview announcement, Ultrafast delivers:
- 14x Speedup: Compared to the standard GPT-5.6 Sol API, which typically hovers around 50-55 tokens per second depending on server load.
- Zero Intelligence Degradation: The model weights and architecture are identical to the standard tier.
- API-First Access: Currently in a limited preview for select enterprise customers, with broader rollout planned as hardware capacity scales.
While competitors like Anthropic have introduced accelerated versions of their models (such as Claude's "fast mode"), they haven't touched the 750 t/s threshold for a true frontier-class model. OpenAI has effectively unbundled speed from model size.
The Hardware Secret: Cerebras and the Wafer-Scale Advantage
You cannot hit 750 tokens per second on a frontier model using standard Nvidia H100s or B200s without running into the memory-bandwidth wall.
Generating text one token at a time is a memory-bound process. For every single token generated, the hardware must move the model's massive parameter weights from memory to the compute units. On conventional GPU clusters, this constant shuttling of data across the PCIe bus or NVLink is the ultimate bottleneck. You can add more compute, but you can't feed the data fast enough.
Enter Cerebras.
OpenAI's Ultrafast mode is powered by a strategic partnership with the AI hardware challenger, utilizing their wafer-scale chips. Instead of stitching together dozens of smaller GPUs, Cerebras builds a single chip the size of a dinner plate. This massive silicon real estate allows for an unprecedented amount of on-chip SRAM memory.
Because the entire GPT-5.6 Sol model (or massive chunks of it) can live directly on the chip, the memory-movement wall is obliterated. The weights don't have to travel across a motherboard; they are already sitting millimeters away from the compute cores. Remove the data-transfer bottleneck, and the exact same model simply emits tokens faster.
OpenAI's decision to productize this hardware advantage is a massive validation of the wafer-scale approach for LLM inference.
Why Speed is Now a Product: The Agentic Multiplier
A single ultra-fast chat reply is a nice luxury. But OpenAI didn't build a Cerebras cluster just so ChatGPT could write a poem a few seconds faster. The real target for Ultrafast is agentic workflows.
In an agentic system, an AI doesn't just answer a prompt; it executes a multi-step plan. It might search a database, read the results, write code, execute the code, read the error log, and rewrite the code.
In these sequential workflows, latency compounds brutally.
- 1 Model Call: Standard mode takes ~9 seconds. Ultrafast takes ~0.7 seconds.
- 10 Sequential Calls: Standard mode takes ~90 seconds. Ultrafast takes ~7 seconds.
- 50 Sequential Calls: Standard mode takes ~8 minutes. Ultrafast takes ~35 seconds.
When an agent takes 8 minutes to complete a task, it feels like a batch job. You submit the request, go get coffee, and come back. When that same task takes 35 seconds, it feels like an interactive tool.
This is why OpenAI is aggressively targeting latency-sensitive enterprise use cases with the Ultrafast preview:
- Live Incident Response: Where cybersecurity agents need to analyze logs and mitigate threats in seconds, not minutes.
- Financial Market Analysis: Where trading algorithms rely on LLMs to parse breaking news and SEC filings instantly.
- Voice Applications: Where even a 500ms delay breaks the illusion of a natural human conversation.
- Customer Support: Where autonomous agents can resolve complex, multi-step database queries while the customer is still typing their next sentence.
The Economic Implications of Speed
Beyond the user experience, there is a fundamental throughput equation at play here. Serving the same workload faster on specialized hardware can drastically alter the cost-per-task for enterprise deployments. While OpenAI has not yet released the official pricing tiers for Ultrafast, the economics of wafer-scale inference suggest a fascinating shift.
Historically, running a massive frontier model required tying up expensive GPU clusters for extended periods. If a Cerebras-powered system can process 14 times the tokens in the same window, the capital expenditure on hardware yields a significantly higher throughput of completed tasks. For high-volume enterprise customers, paying a premium for Ultrafast might actually result in a lower total cost of ownership when factoring in the time saved by human operators and downstream systems waiting on the AI's output.
Furthermore, this puts immense pressure on cloud providers and competing AI labs. If developers get used to 750 tokens per second as the new baseline for agentic workflows, the standard 50 tokens per second will quickly feel obsolete. We can expect to see a massive scramble among competitors to secure their own specialized inference hardware, moving the battleground away from standard Nvidia clusters and toward custom silicon solutions designed explicitly to break the memory-bandwidth wall.
The Inference Speed War Enters a New Phase
For the last two years, the AI industry has competed on two main axes: benchmark intelligence (Elo scores) and price per million tokens.
With the launch of Ultrafast, OpenAI has formally introduced a third axis: premium inference speed.
By turning hardware-accelerated speed into a purchasable API tier, OpenAI is changing the calculus for developers. You no longer have to spend weeks fine-tuning an 8B parameter model just to get your app's latency down to acceptable levels. If the unit economics make sense, you can simply pay for Ultrafast and get GPT-5.6 Sol intelligence at small-model speeds.
The preview is currently limited, and OpenAI has not yet revealed the exact pricing multiplier for Ultrafast tokens. But the writing is on the wall: the future of AI isn't just about who has the smartest model. It's about who can serve that intelligence fast enough to make autonomous agents actually usable in the real world. And right now, with Cerebras under the hood, OpenAI is setting the pace.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.