Technological Advances in AIGenerative AI

GPT-5.6 Sol Accelerates 10x: OpenAI Unveils Its New Ultrafast Mode

A few weeks after rolling out GPT-5.6 on a large scale, OpenAI is tackling another major challenge in artificial intelligence: latency. With its new Ultrafast mode, the company aims to enable its flagship model, GPT-5.6 Sol, to generate responses at a speed previously associated mainly with much lighter models. The advertised throughput can reach 750 tokens per second, while retaining the same model and thus, according to OpenAI, without sacrificing its reasoning capabilities.¹ This acceleration might seem purely technical, but it actually addresses a key challenge in agent-based AI. When an agent must analyze information, call upon multiple tools, verify a result, and then continue its work, every second of latency compounds as the process progresses. With Ultrafast, OpenAI specifically aims to reduce this wait time and bring AI agents closer to true real-time execution.

GPT-5.6 Sol is the most powerful model in the GPT-5.6 family, ahead of Terra and Luna, and is designed for complex tasks in programming, research, cybersecurity, science, and automation. OpenAI had previously announced that it runs on Cerebras’ infrastructure at speeds of up to 750 tokens per second—a level far exceeding that of traditional inference.¹ The appeal of Ultrafast lies primarily in the fact that OpenAI does not present this acceleration as a new, streamlined version of the model. The concept involves running GPT-5.6 Sol on infrastructure specifically optimized for ultra-low-latency inference, in order to maintain its level of intelligence while significantly reducing the time required for generation.

The difference from traditional acceleration methods is significant. Until now, achieving a faster response often meant choosing a smaller model or reducing the reasoning effort. Ultrafast, on the other hand, seeks to further decouple intelligence from speed by optimizing the infrastructure that runs the model. In practice, reaching 750 tokens per second means that long responses can appear almost instantly once generation begins. The actual result, however, depends on many factors, including initial reasoning time, context length, calls to external tools, and infrastructure load. OpenAI also notes that actual performance may vary depending on workloads.¹

MSc in Generative and Agent-Based AI at aivancity
aivancity

MSc in Generative,
, and Agent-Based AI

Become an expert in generative and agent-based AI: LLMs, transformers, and enterprise deployment. Includes a learning trip to Silicon Valley. RNCP Level 7 certification (equivalent to a 6-year post-secondary degree).

12 months — 6 years of post-secondary education Master's degree (Bac+5) with experience Part-time — Fridays & Saturdays Paris-Villejuif Campus
Learn more about the program → RNCP Level 7 Certification

To achieve this level of performance, OpenAI relies on Cerebras Systems, a U.S. company specializing in processors designed for artificial intelligence. Its approach differs from that of traditional GPU infrastructure: Cerebras develops processors on the scale of an entire silicon wafer—known as Wafer-Scale Engines—to bring memory and computing units closer together and reduce data transfers that slow down inference.²

This collaboration goes far beyond the launch of a simple fast mode. In 2026, OpenAI and Cerebras entered into a multi-year agreement to deploy 750 MW of inference capacity, with a gradual ramp-up of the infrastructure through 2028.² The exact amount associated with the partnership has varied across announcements and financial documents, but it runs into the tens of billions of dollars, demonstrating the strategic importance OpenAI places on high-speed inference.³ Behind Ultrafast, a more profound shift in the market is taking shape: after several years spent training increasingly powerful models, research labs are now seeking to optimize how these models operate once deployed.

In a typical conversation with ChatGPT, saving a few seconds mainly makes the experience more comfortable. For an agent-based AI, the impact can be much greater. An autonomous agent doesn’t necessarily generate just one response. It can search for information, analyze a file, query a database, execute code, verify the result, and then repeat the process until it achieves its goal. Each step engages the model and adds latency to the overall workflow.

With much faster inference, these workflows can become significantly smoother. A developer could task an agent with analyzing logs and searching for an error while another prepares a fix. A financial analyst could analyze multiple sources simultaneously before generating a summary. In customer support or voice interfaces, reduced latency could make interactions with artificial intelligence feel much more natural. OpenAI is already positioning GPT-5.6 Sol as a particularly high-performing model for agent-based tasks and software development, notably thanks to its results on Terminal-Bench, coding agent evaluations, and long-term professional workflows.⁴

However, speed does not eliminate all bottlenecks. An agent remains dependent on the time required to access an API, load a document, perform a search, or obtain human validation. Ultrafast accelerates the reasoning engine, but it does not instantly speed up the entire IT system surrounding it. That is why the benefits will likely be most noticeable in applications where inference currently accounts for a significant portion of the total execution time.

Executive MBA in AI & Business Transformation, aivancity
aivancity

Executive MBA in AI & Business Transformation

The MBA Redesigned for the Age of AI. For experienced executives who want to lead the transformation of their organizations. Paris, Nice, and Dubai.

12 months — Part-time At least 10 years of experience Early bird: 20,000 € Paris · Nice · Dubai

This is where Ultrafast’s promise reaches its main limit. An infrastructure capable of running a boundary model at several hundred tokens per second requires considerable hardware capacity. When GPT-5.6 was launched, OpenAI had already indicated that access to Sol execution on Cerebras would initially be limited to certain customers, while the available capacity was gradually increased.¹

This situation stands in contrast to OpenAI’s broader strategy surrounding GPT-5.6. The family is now available on a much wider scale, with Sol as the flagship model, Terra to balance performance and cost, and Luna as a faster and more cost-effective option. GPT-5.6 Sol is priced at $5 per million input tokens and $30 per million output tokens via the API, while Terra costs $2.50 and $15, respectively, and Luna costs $1 and $6.⁴ Access to certain advanced capabilities also depends on the ChatGPT or Codex plan used. For example, OpenAI is gradually rolling out GPT-5.6 Sol to eligible plans, while the most advanced reasoning options are reserved for certain subscriptions.⁵

For Ultrafast, the real challenge will therefore be as much economic as it is technological. If the cost of this acceleration is high, it will be most relevant for applications where every second has direct value: real-time assistance, interactive software development, incident response, chatbots, transactions, or operational analysis. For background processes, such as document indexing or overnight analyses, a less expensive standard mode might remain a much more practical option.

The partnership with Cerebras also signals a shift in the dynamics of competition among artificial intelligence players. Until now, the race was primarily measured by the size of models, their benchmark performance, or their ability to reason. Now, the time it takes to achieve this level of intelligence is itself becoming a competitive advantage.

For OpenAI, working with Cerebras also allows it to diversify its computing infrastructure in a market historically dominated by Nvidia GPUs. Cerebras claims that its specialized architectures help reduce certain bottlenecks related to data movement and achieve particularly high throughput for inference.² OpenAI is not replacing its existing infrastructure, but is adding a new category of resources optimized for workloads where latency is critical.

This trend could gradually change the way companies choose their models. The best AI will no longer necessarily be the one that achieves the highest absolute score on a benchmark, but rather the one that offers the best balance between quality, cost, token consumption, and execution time.

Making artificial intelligence ten times faster is not just a technical advancement. The faster agents can analyze, decide, and act, the more important the issue of human oversight becomes. In a traditional agent-based system, latency still allows time to verify an action, interrupt a process, or request validation. When multiple agents can chain together decisions in a matter of seconds, this window for intervention can shrink considerably.

This issue is particularly sensitive in cybersecurity, finance, critical infrastructure, and automated administrative processes. GPT-5.6 Sol already has advanced cybersecurity capabilities, and OpenAI acknowledges that evaluations cannot cover all possible combinations of models, tools, and workflows. The company states that it has strengthened its security mechanisms with real-time controls, monitoring, human and automated red teaming, and risk-based access controls.¹

A second challenge concerns unequal access to computing power. If the fastest versions of the best models remain available only to large companies that can afford premium infrastructure, a gap could widen between organizations with near-instantaneous agents and those using slower systems. Speed would then become an economic advantage in its own right, just like access to the most powerful models.

Finally, this acceleration raises an environmental concern. The infrastructure needed to run state-of-the-art models at very high speeds consumes a considerable amount of electricity. The 750 MW agreement between OpenAI and Cerebras illustrates the industrial scale now required to operate modern AI systems.² The race for lower latency will therefore need to be evaluated not only in terms of tokens per second, but also in terms of energy efficiency and the actual value produced by each computation.

With Ultrafast, OpenAI demonstrates that the next frontier in artificial intelligence is no longer just about making models smarter. It’s also about making that intelligence fast enough to keep pace with human activities and computer systems. GPT-5.6 Sol can already reason, code, use tools, and coordinate complex workflows. By achieving up to 750 tokens per second on the Cerebras infrastructure, these capabilities can now be harnessed with significantly reduced latency.¹

This development could be particularly significant for AI agents, whose effectiveness depends as much on their reasoning ability as on the time required to complete each step. But Ultrafast also highlights new tensions within the industry: infrastructure costs, unequal access to premium capabilities, energy consumption, and the need to maintain human oversight as automated systems act ever more quickly. The next competition between OpenAI, Anthropic, Google, and other labs may therefore no longer focus solely on the question “Which AI is the smartest?” but also on another, increasingly critical one: How long must we wait for it to act?

Technology Framework

How does GPT-5.6 Sol's Ultrafast mode work?

Ultrafast mode is a new execution option developed by OpenAI for GPT-5.6 Sol, with the goal of significantly reducing latency without replacing the model with a lighter version. Unlike approaches that prioritize speed by reducing reasoning capabilities, Ultrafast retains GPT-5.6 Sol while optimizing the infrastructure used for its inference. The system can thus generate up to 750 tokens per second, which significantly speeds up the production of long responses and the execution of tasks requiring multiple successive interactions with the model.

This acceleration relies in particular on the infrastructure of Cerebras Systems, which specializes in processors designed for artificial intelligence. Its Wafer-Scale Engine architecture utilizes a very large silicon die combined with memory placed as close as possible to the computing units. This design aims to minimize the constant data transfers between the processor and external memory, which are one of the main factors slowing down the inference of large language models.

Ultrafast plays a particularly important role in the field ofagent-based AI. An agent does not necessarily produce a single response; it can analyze a request, consult multiple tools, execute code, examine a result, correct an error, and then continue its work. Each step potentially requires a new call to the model. By accelerating GPT-5.6 Sol, OpenAI is therefore seeking to reduce the duration of these successive loops and bring certain agent-based workflows closer to real-time execution.

This approach could be particularly useful for software development, data analysis, conversational support, financial research, and incident response. However, the model’s speed does not eliminate other sources of latency. Calls to external APIs, database access, document loading, and human validation can still slow down the overall workflow. Ultrafast therefore primarily accelerates the part of the process directly related to GPT-5.6 Sol inference.

Key Features of Ultrafast Mode
  • Ultra-high-speed inference: generation of up to 750 tokens per second, depending on usage conditions
  • GPT-5.6 Saved: Acceleration of the flagship model without any announced plans to use a lighter version
  • Cerebras Infrastructure: Using a Specialized Hardware Architecture to Reduce Data Movement and Latency
  • Optimizing AI Agents: Accelerating Reasoning, Tool Use, Verification, and Execution Loops
  • Software development: the ability to make workflows for code generation, analysis, and correction more interactive
  • Real-time applications: useful for voice assistants, customer support, operational analytics, and incident response
  • API Integration: A mode designed for professional use and for applications that leverage GPT-5.6 Sol through the OpenAI ecosystem
Technical constraints and limitations
  • Ultrafast still relies on specialized and significant computing capabilities, which may limit its large-scale deployment
  • Access remains restricted, and its opening depends on the availability of infrastructure capacity
  • The advertised maximum data rate does not necessarily correspond to the actual speed achieved in all situations
  • Calls to tools, APIs, databases, or external systems may continue to slow down agent workflows
  • An increase in processing speed does not automatically guarantee improved accuracy or reliability of the responses
  • Systems that operate more quickly require appropriate monitoring mechanisms, particularly when agents can perform several successive autonomous actions
  • The cost and availability of large-scale infrastructure are key factors in the adoption of this approach by businesses

With its Ultrafast mode, GPT-5.6 Sol exemplifies the search for a new balance between reasoning power, execution speed, and user-friendliness in artificial intelligence models. On a related topic, check out our article “OpenAI Launches GPT-5.5 Instant with Fewer Errors and Faster Responses, ” which analyzes the previous step in this evolution toward models capable of responding more quickly while improving the reliability of their results.

1. OpenAI. (2026). Previewing GPT-5.6 Sol: A Next-Generation Model.
https://openai.com/index/previewing-gpt-5-6-sol/

2. Cerebras Systems. (2026). OpenAI Partners with Cerebras to Bring High-Speed Inference to the Mainstream.
https://www.cerebras.ai/blog/openai-partners-with-cerebras-to-bring-high-speed-inference-to-the-mainstream

3. Reuters. (2026). OpenAI Signs $10 Billion Computing Deal with Nvidia Challenger Cerebras.
https://www.reuters.com/

4. OpenAI. (2026). GPT-5.6: Frontier Intelligence That Scales with Your Ambition.
https://openai.com/index/gpt-5-6/

5. OpenAI Help Center. (2026). GPT-5.6 in ChatGPT.
https://help.openai.com/en/articles/20001354-gpt-56-in-chatgpt/

Don't miss our upcoming articles!

Get the latest articles written by aivancity experts and professors delivered straight to your inbox.

We don't send spam! Please see our privacy policy for more information.

Don't miss our upcoming articles!

Get the latest articles written by aivancity experts and professors delivered straight to your inbox.

We don't send spam! Please see our privacy policy for more information.

Related posts
Technological Advances in AIAI & Robotics

Skild AI's S1: How Video Learning Is Truly Changing Robotics

Unveiled by Skild AI in August 2026, S1 is a foundational model for robotics designed to perform a new task based on a single video demonstration, without the need for additional task-specific training. The…
AI & EducationGenerative AI

ChatGPT for Teens: OpenAI Wants to Guide the Generation Growing Up with AI

On August 18, 2026, OpenAI launched a ChatGPT pilot program for 13- to 17-year-olds, featuring enhanced safeguards, age estimation, learning tools, and parental controls. The program is not a…
Generative AI

Gemini 3.7 Flash: Google Boosts Its AI for Coding and Automation

Just three weeks after Gemini 3.6 Flash, Google is already stepping up the pace with Gemini 3.7 Flash, a new model designed to boost performance in programming, automation, and workflows involving multiple AI agents…