Newsroom
13 August, 2026 / News / AI / Tags: flash, gpt, gemini, ultrafast, google

Google makes its latest coding and agent model widely available at low cost, while OpenAI tests a Cerebras-powered speed tier limited to select users
Google has released Gemini 3.7 Flash as a broadly accessible artificial intelligence model aimed at software engineering, web development and autonomous agents. At the same time, OpenAI has opened a restricted preview of GPT-5.6 Sol Ultrafast, a faster service tier of its existing top model that relies on specialized hardware from Cerebras.
The announcements, made on August 13, 2026, underscore a shift in the industry toward models that deliver rapid responses suitable for real-time interactive applications rather than solely competing on raw capability scores.
Gemini 3.7 Flash supports inputs of up to one million tokens, equivalent to roughly 750,000 words, and processes multiple formats including text, images, video, audio and PDFs. It includes tool-calling functions that allow it to control computers and manage multi-step workflows.
Google reports that the model completes a standard coding task in two minutes and 13 seconds, less than half the time required by the previous Flash version, while also improving output quality. The company positions it as an efficient workhorse for developers and enterprise users who need high-volume, low-latency results.
The model is available immediately to Google Cloud developers in more than 160 countries. Introductory pricing stands at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. After that date the rates rise to $1.50 and $7.50 respectively.
According to Google’s own benchmarks, Gemini 3.7 Flash ranks ahead of Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 test categories. It records a Code Arena web-development Elo score of 1,588 and achieves 30.4 percent on the AutomationBench enterprise workflow metric.
OpenAI’s new Ultrafast tier is not a separate foundation model. It runs the existing GPT-5.6 Sol architecture on Cerebras wafer-scale chips, delivering throughput of up to 750 tokens per second, or about 560 words per second. The company states this represents roughly 14 times the standard speed of GPT-5.6 Sol.
The higher throughput is intended to support applications such as live voice agents that must process information and respond during ongoing conversations. GPT-5.6 Sol itself had previously undergone additional security hardening against prompt-injection attacks.
Access remains limited to an invited group of business customers through the OpenAI API. Broader availability is planned as infrastructure capacity expands.
Google has chosen wide geographic distribution and aggressive introductory pricing to place Gemini 3.7 Flash in the hands of a large developer base. OpenAI has prioritized peak performance for a smaller set of users while depending on third-party hardware rather than proprietary chips for the speed gains.
Both companies are responding to rising demand for AI systems that can operate as autonomous agents in real time. Gemini 3.7 Flash is already live for general use, while GPT-5.6 Sol Ultrafast remains in limited preview.









