OpenAI Unveils GPT-5.6 Sol Ultrafast With Speeds of Up to 750 Tokens per Second

16.08.2026 3 minutes Author: Newsman

OpenAI has introduced a new high-speed mode, GPT-5.6 Sol Ultrafast, which runs on Cerebras infrastructure and can generate up to 750 output tokens per second. That makes it roughly 14 times faster than the model’s standard mode and significantly reduces the time required to complete large tasks.

According to OpenAI, the new mode is designed primarily for scenarios where response speed directly affects the workflow. Instead of waiting for a lengthy task to finish, users can receive results almost in real time and immediately move on to the next iteration.

The difference is particularly noticeable when working with large amounts of text or code. For example, generating a code patch of around 2,000 tokens at a speed of 750 tokens per second takes approximately 2.7 seconds.

If that speed is maintained throughout the entire generation process, even the maximum output limit of 128,000 tokens could theoretically be reached in under three minutes. For complex tasks where the model needs to produce large amounts of code, analysis, or structured data, this could significantly change the way users interact with AI.

OpenAI describes the advantage of Ultrafast as enabling “more useful work per second.” The company believes the increased speed could turn some lengthy processes into interactive ones.

For example, tasks that previously had to be left running overnight because of long processing times could potentially be completed during the workday, allowing users to review results and make changes immediately. This makes it possible to perform more iterations without having to wait a long time for each new response from the model.

Ultrafast could be particularly useful for programming, working with large codebases, data analysis, and other tasks where the model needs to generate substantial amounts of information. Its high speed also makes interacting with the model feel more like working with a real-time interactive tool.

Importantly, this is an accelerated mode of GPT-5.6 Sol rather than a separate model with fundamentally different capabilities. The main difference is inference speed, enabled by specialized Cerebras infrastructure.

For now, GPT-5.6 Sol Ultrafast is not available to everyone. OpenAI has launched the technology in limited preview and is providing access only to a select group of customers.

The company plans to gradually expand availability as additional compute capacity comes online. OpenAI has already opened a dedicated waitlist form for users and companies interested in gaining access to Ultrafast in the future.

Ultimately, OpenAI’s focus with Ultrafast is not on increasing the size of the model, but on dramatically reducing the time between a request and the resulting output. If the technology becomes widely available, complex tasks involving large amounts of generated content could shift from lengthy background processing to an almost fully interactive workflow.

Subscribe
Notify of
0 Коментарі
Oldest
Newest Most Voted
Found an error?
If you find an error, take a screenshot and send it to the bot.