On Wednesday, Google shipped Gemini 3.7 Flash, a lightweight multimodal model that is faster, cheaper, and more obedient to developer instructions than its predecessor. The announcement, posted to the company’s blog and picked up by Thurrott.com, was brisk and businesslike — a routine update in a product line that now seems to refresh every few weeks.

It is also, if you listen to the right people, a sign of catastrophic failure.

The catastrophe in question is Gemini 3.5 Pro, Google’s would-be flagship. First teased months ago, the model has missed at least two public deadlines. Forbes reported on Wednesday that some are calling it “the longest-awaited model of 2026.” Rumors now swirl of senior researcher departures, persistent coding failures, and a possible complete retraining from scratch — a “structural problem” so deep the entire pre-training run may need to be redone. The subtext of every headline is the same: Google is losing the AI race, and losing badly.

That is one way to read the situation. Here is another: the industry’s fixation on flagship models is a marketing hangover, and Google’s rapid Flash cadence — 3.6 three weeks ago, 3.7 now — reveals a company that has quietly figured out where the economics of AI actually point.

The Pro That Never Was

Let’s state the obvious. A delayed flagship is embarrassing. Developers who planned roadmaps around 3.5 Pro are annoyed. Investors who want a narrative of American technological dominance are nervous. And Google’s communications team, which has spent two years promising a model that would “set a new standard,” now has to explain why the standard is being set by a series of smaller, cheaper models nobody asked to be excited about.

But the delay itself is worth examining. According to the Forbes report, the problems are not cosmetic. They involve “persistent coding and reliability problems” — the kind of issues that, in a less breathless era, would be called “the product doesn’t work yet.” A company that ships a broken flagship does not win the AI race. It wins a news cycle and then a reputation for unreliability. Ask any enterprise buyer who got burned by an overhyped model launch in 2024 how eager they are to repeat the experience.

One engineer at a mid-sized fintech firm, reached via Slack while his team was evaluating the new Flash endpoint, put it bluntly: “I don’t care if the Pro model is late. I care if the model I’m actually paying for hallucinates on a compliance query. Flash 3.7 is cheaper than 3.6 and follows instructions better. That’s my entire decision criteria.”

Flash Forward

The numbers, such as they are, tell a story the Pro-obsessed coverage misses. Gemini 3.7 Flash, per Google’s announcement, offers a lower cost per task on average than its predecessor. It is better at following developer instructions. It is multimodal. It is available now, on Vertex AI and Gemini Enterprise, for anyone who wants to build on it.

Three weeks ago, the company shipped Gemini 3.6 Flash. Before that, a steady cadence of Flash releases stretching back through the year. Each one is incrementally better, incrementally cheaper, and incrementally more useful for the high-volume, agent-driven workloads that actually consume inference compute. This is not the behavior of a company in disarray. It is the behavior of a company that has decided the real game is commoditizing inference, and it intends to be the low-cost provider.

That strategy has a name in other industries: it is called winning. Walmart did not become Walmart by releasing a single, perfect flagship store. It became Walmart by making the economics of retail work at scale, iterating relentlessly on logistics while competitors obsessed over flagship locations in Manhattan.

The Economics Nobody Talks About

The AI industry’s discourse runs on a single metric: benchmark scores. Which model tops the leaderboard? Which model aces the graduate-level reasoning test? The assumption, rarely examined, is that the model with the highest score is the one that matters.

But inference is not a benchmark. Inference is a cost center. Every enterprise deploying AI at scale — customer service, document processing, code generation, agentic workflows — is making a procurement decision, not a prestige decision. They are asking: what does it cost per million tokens, how fast is the response, and does it follow instructions reliably enough to automate the task? On all three questions, the Flash line is designed to win.

Google’s critics will counter that the company is being forced into this position — that it cannot ship a competitive frontier model, so it is dressing up its consolation prize as a strategy. The counter-counter is that the frontier model business is a terrible business. Training costs are astronomical. Inference costs are higher. Margins are thin or negative. The customers who need a model that can solve Olympiad-level math problems are a rounding error compared to the customers who need a model that can summarize 10,000 support tickets for a fraction of a cent each.

If Google eventually ships Gemini 3.5 Pro and it is excellent, the Flash strategy will look prescient — a two-tier offering that covers both prestige and volume. If 3.5 Pro never ships, or ships and disappoints, the Flash line will still be there, quietly handling the workloads that actually pay the bills. Either way, the company that iterates fastest on cheap inference has a durable advantage. Right now, that company is Google.

None of this is to excuse the communications mess or the missed deadlines. Google promised a flagship and has not delivered it. That is a failure of expectation-setting, and the people who set those expectations should answer for it. But the product that actually shipped this week — the one you can use, right now, for less money than last month’s version — is not a failure. It is a signal, if you are willing to read it, about what the AI market is actually going to value.

Sources