DMR News

Advancing Digital Conversations

OpenAI Launches Ultrafast Mode for GPT-5.6 Sol at Up to 750 Tokens per Second

ByJolyen

Aug 16, 2026

OpenAI Launches Ultrafast Mode for GPT-5.6 Sol at Up to 750 Tokens per Second

OpenAI has introduced Ultrafast, a new API service tier that can run GPT-5.6 Sol up to 14 times faster than standard processing. The company says the system can generate as many as 750 output tokens per second while using the same underlying model.

Ultrafast is launching first through the OpenAI API in a limited preview. OpenAI said access will initially be restricted to a small group of customers and will expand as additional capacity becomes available.

In its official announcement, OpenAI said the service is designed for applications where response speed is critical without requiring developers to switch to a smaller or more specialized model. The company described the goal as increasing the amount of useful work a model can complete each second.

Cerebras Powers the New Service Tier

Ultrafast is powered by OpenAI’s partnership with AI chipmaker Cerebras. OpenAI says the infrastructure allows GPT-5.6 Sol to reach output speeds of up to 750 tokens per second, compared with its normal API processing tier.

The faster processing could be useful for workloads that require frequent or time-sensitive model responses. OpenAI highlighted examples including incident response, customer support, financial market analysis, and e-commerce applications.

The service does not introduce a separate model with reduced capabilities. Instead, OpenAI is using a different inference configuration and hardware setup to run GPT-5.6 Sol at substantially higher speeds.

AI Labs Are Competing on Inference Speed

Anthropic offers a similar option through Claude Fast Mode, which runs supported Claude Opus models at up to 2.5 times their standard output speed. Anthropic says Fast Mode uses the same model weights and capabilities while charging higher API prices for the faster inference tier.

OpenAI’s reported 14-times improvement is considerably larger, although the two companies measure performance against their own standard processing systems and support different models and infrastructure. Anthropic’s Fast Mode is also currently limited to supported Claude Opus models through its API.

OpenAI has not announced general availability or broader pricing details for Ultrafast. For now, the feature remains an early preview intended to test demand and performance before wider access.


Featured image credits: Focal Foto via Flickr

For more stories like it, click the +Follow button at the top of this page to follow us.

Jolyen

As a news editor, I bring stories to life through clear, impactful, and authentic writing. I believe every brand has something worth sharing. My job is to make sure it’s heard. With an eye for detail and a heart for storytelling, I shape messages that truly connect.

Leave a Reply

Your email address will not be published. Required fields are marked *