On Wednesday, Google released Gemini 3.7 Flash — a refinement of the 3.6 Flash model it shipped exactly three weeks earlier. The pricing is eye-catching: $0.75 per million input tokens, $3.75 per million output tokens, a rate that undercuts Anthropic’s Claude Sonnet 5 and OpenAI’s GPT-5.6 Terra by roughly 60 percent. The introductory price expires December 31, at which point it doubles.

Meanwhile, Gemini 3.5 Pro — the flagship model Wall Street has been waiting for since Sundar Pichai promised it at I/O in May — remains missing. Its original June target came and went. Internal deadlines have slipped repeatedly. Reuters reported in July that the model was falling short of internal goals, particularly on coding benchmarks. As of this morning, 67 days past the original window, there is no launch date.

The conventional read on this split-screen is straightforward: Google can ship the small stuff but can’t compete at the high end. The Flash updates are consolation prizes while DeepMind struggles to get the real model out the door. It’s a narrative that writes itself, and it’s probably wrong.

The Numbers That Actually Matter

Look at what 3.7 Flash delivers. FrontierCode jumps from 34.4 percent to 43.6 percent. DeepSWE climbs from 48.6 percent to 65.3 percent. On the WebDev Arena leaderboard, it posts a 1588 Elo. These are not cosmetic bumps — they are material gains on the benchmarks that enterprise buyers actually care about, shipped on a three-week cycle.

Now look at the pricing structure. The $0.75 rate is explicitly introductory. Google is not giving away margin out of charity; it is setting a hook. Developers who build on 3.7 Flash at the discount rate will face a doubling of their inference costs in January, by which point their pipelines, prompts, and fine-tuned workflows will be built around the Flash architecture. Switching costs will have accrued. The introductory window is long enough to build dependency and short enough to create urgency.

This is not the behavior of a company that views Flash as a stopgap while it figures out Pro. It is the behavior of a company that believes Flash is the product.

The Tier That Ate the Stack

There is a pattern in enterprise software that repeats often enough to be a law: the “good enough” tier eventually consumes the premium tier from below. It happened with cloud compute, where AWS’s commodity instances eroded the market for specialized hardware. It happened with SaaS, where lightweight tools ate into the consulting- and customization-heavy deployments of the previous era. It is happening now with AI models, and Google appears to be the only lab that has internalized the lesson.

Anthropic and OpenAI are fighting a benchmark war at the frontier, where each percentage-point gain on graduate-level reasoning tests costs tens of millions in compute and generates headlines but marginal revenue. Google, by contrast, is iterating on the model that developers actually put into production — the one where token costs show up on a monthly invoice. The Flash line is not the junior varsity squad waiting to be promoted. It is the main event.

A developer at a mid-sized fintech firm, reached via Slack while his team was evaluating model providers for a customer-support pipeline, put it plainly: “We looked at Sonnet and Terra. The benchmarks are better. But when you run the numbers on a million API calls a day, the difference in output quality doesn’t justify a 5x cost multiplier. Flash is already past the threshold where the errors we see are the kind a human would make anyway.” He paused. “I don’t think we’re going to bother testing Pro when it ships.”

What Wall Street Is Missing

The analyst class has fixated on Gemini 3.5 Pro as the signal of Google’s AI competitiveness. Every missed deadline is treated as a data point in a story about DeepMind falling behind. But the fixation on the flagship obscures a more consequential question: what if the market for flagship models is smaller than everyone assumes?

Enterprise AI adoption does not turn on whether a model can solve International Math Olympiad problems. It turns on whether the model is reliable enough to handle customer email, cheap enough to deploy at scale, and fast enough that users don’t tab away while waiting for a response. Flash clears all three bars, and it is improving faster than the frontier models are pulling away.

If that trend holds — and three weeks between meaningful updates suggests it will — then the delay of 3.5 Pro is not a crisis for Google. It is an accounting footnote. The product that matters already shipped. Twice this summer.

The real risk for Google is not that Pro is late. It’s that by the time Pro arrives, nobody will need it.

Sources