On Tuesday, Google DeepMind released three new AI models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. It also killed one.
Buried in the announcement was the news that Gemini 3.5 Flash — the “workhorse” model that headlined Google I/O in May, the one developers were told to build on, the one pitched as the future of cost-efficient inference — has already been deprecated. Two months. That’s how long the star of the show lasted before Google yanked it offstage and told everyone to move to the next thing.
The tech press dutifully covered the new models. 3.6 Flash uses 17% fewer output tokens. Flash-Lite costs $0.30 per million input tokens. Flash Cyber finds security vulnerabilities. All useful, all incrementally better. But the coverage missed the story that will actually matter to anyone running a business on these APIs: the churn is now the product.
The Two-Month Model
Deprecation in enterprise software usually comes with migration guides, long sunset periods, and enough lead time for a CFO to budget the transition. Google just gave the market 60 days. Not 60 days’ notice — 60 days of total existence for a model that was supposed to be foundational.
This is not a one-off. The AI industry has normalized a release cadence that would be considered reckless in any other part of the enterprise stack. Imagine AWS deprecating an EC2 instance type two months after launch and telling customers to rewrite their provisioning scripts. There would be revolt. But in AI, the assumption is that faster is always better, and anyone who can’t keep up simply lacks vision.
A software architect at a mid-sized logistics firm put it plainly in a private Slack channel for enterprise AI practitioners: “We finished our migration to 3.5 Flash three weeks ago. Our legal team signed off on the data handling addendum last Tuesday. Now it’s deprecated. My CEO wants to know if we’re bad at planning or if Google is.”
The answer, increasingly, is both.
The Planning Tax
Google is not alone here. OpenAI, Anthropic, and the rest have all trained the market to expect constant model turnover. But Google’s move is unusually aggressive because it couples rapid deprecation with a pricing structure that looks cheap on a per-token basis but imposes a hidden planning tax on every organization that adopts these models in production.
That tax takes real forms: re-tuning prompts, re-running evaluation suites, re-negotiating compliance reviews, re-training internal documentation. For a regulated industry — healthcare, financial services, anything touching government contracts — each model swap can trigger a mini-audit. The per-token savings vanish the moment you factor in the engineering hours burned on migration.
Google knows this. The 17% token reduction in 3.6 Flash is a genuine efficiency gain. But it’s also a convenient way to frame forced obsolescence as a gift. “We made it cheaper” sounds better than “we made the old one stop working.”
What Pro’s Absence Really Means
Then there is the missing model. Gemini 3.5 Pro was promised for June, then pushed to July 17, and as of this week it still isn’t here. Four senior Gemini researchers decamped to Anthropic earlier this year, according to reporting from Bind AI. The Pro delay is now on its third rescheduling, with leaks citing persistent hallucination problems and underwhelming benchmark performance against GPT-5.6.
Read those two facts together — the frantic release of small Flash variants and the continued absence of a flagship Pro model — and a different picture emerges. Google is not iterating because it has mastered the art of rapid improvement. It is iterating because it cannot ship the thing it actually needs to ship, and it has to keep the developer ecosystem fed with something in the meantime.
Flash-Lite and Flash Cyber are real products with real use cases. But they are also placeholders. They fill a gap in the roadmap while the Pro team struggles to close the gap with competitors who are not standing still.
None of this means the new models are bad. They aren’t. The benchmarks are solid, the pricing is aggressive, and the cybersecurity specialization is genuinely clever. But the story of this release is not the technology. It’s the tempo. And the tempo is starting to look less like innovation and more like panic.
Businesses that bet on any single model today are making a wager not just on its capabilities, but on its shelf life. Right now, that shelf life is measured in weeks. That is not a foundation to build on. It’s a treadmill, and Google just turned up the speed.
Sources
- Google releases three new Gemini models — but no 3.5 Pro
- Google reveals faster and cheaper Gemini 3.6 Flash, says …
- Gemini 3.6 Flash: Price, Benchmarks, API & Flash-Lite
- Google Delays Gemini 3.5 Pro to July 17
- Gemini 3.5 Pro Delay: Why Google Postponed Its AI Again
- Gemini 3.5 Pro Delayed to July 2026 - Bind AI Blog