People are misreading the Gemini 3.7 and 3.8 Flash releases. Google just showed it can post-train again, which is why I expect Gemini 4 to be good
Peter Van Dijck · September 3, 2026 · 8 min read
My take: three Flash models in six weeks means Google has a post-training loop that works. That might also mean that Gemini 4 will be much better than we expect.
Let’s go through it.
Everyone has decided Google is done
Google shipped Gemini 3.8 Flash yesterday, three weeks after 3.7 Flash, which came three weeks after 3.6 Flash. There has been no Pro model since February, and the 3.5 Pro promised for June never arrived. Jeremie Harris on Last Week in AI #253, back when 3.6 Flash came out:
When we’re thinking about the best models in the world, it’s Anthropic and it’s OpenAI, and there’s just not really anyone else.
— Jeremie Harris, Last Week in AI, 29 July 2026
Then SemiAnalysis, on 7 August, in a piece by Dylan Patel and colleagues called “Gemini is Cooked but GCP is Cooking”:
For all intents and purposes, we believe DeepMind is no longer a frontier lab. […] Along with Noam Shazeer and John Jumper, most of the best RL people at Gemini recently left the company.
— SemiAnalysis, 7 August 2026
They ranked Gemini “8th or 9th” and “don’t see Gemini 4 reversing their downfall.” Jordi Hays read the piece out on TBPN that afternoon; John Coogan’s review was “Brutal, brutal stuff.” The departures, Jeff Dean, Sanjay Ghemawat, Quoc Le and Oriol Vinyals, plus Demis Hassabis stepping back from running DeepMind, had been Alex Volkov’s lead story on ThursdAI the day before. His co-host Wolfram Ravenwolf:
Maybe that is why we don’t have the new Gemini model now, because the people were already on their way out.
— Wolfram Ravenwolf, ThursdAI, 6 August 2026
Harris agreed with SemiAnalysis on #254: “look at the path they’re charting. It feels a lot more like the IBM trajectory, unfortunately.”
A week later 3.7 Flash landed and got a shrug. Nathaniel Whittemore on The AI Daily Brief:
While it is neither the much-delayed Gemini 3.5 Pro nor the now increasingly anticipated Gemini 4 […] by optimizing for speed, the model kind of ends up in a strange no man’s land.
— Nathaniel Whittemore, The AI Daily Brief, 14 August 2026
Alexander Wissner-Gross on Moonshots with Peter Diamandis, on why search pressure produces Flash models and no Pro:
Google’s still out of the running for the capability frontier. […] Note that there’s no Gemini 3.7 Pro anywhere. It’s just Flash. It’s small, it’s fast, and it’s reliable. I think this is over-optimized for clock speed.
— Alexander Wissner-Gross, Moonshots, 27 August 2026
Emad Mostaque, same episode: “Google did do a preview of Gemini 3.5 Pro, but it just couldn’t keep up.” Harris on Last Week in AI #255:
They’ve shown they can ship fast, but very conspicuously, the thing that they’re shipping fast is not an actual frontier model. […] I think it’s starting to raise some real questions and doubts about whether Google has the chops to be a frontier lab.
— Jeremie Harris, Last Week in AI, 31 August 2026
Even people who like the models find the cadence a nuisance. On The Cloud Pod (28 August) the hosts joked that at least “you don’t really have to worry too much about migrating off of 3.6,” and that if 3.8 came two weeks later people would riot. It came five days later. Ryan Whitwam at Ars Technica:
Today, Google is announcing its third Flash model release in just six weeks, making it more likely that we’ll never see the promised Gemini 3.5 Pro.
— Ryan Whitwam, Ars Technica, 2 September 2026
So the story: the RL people left, the Pro is stuck, and Google ships small models because small models are all it can ship.
Three Flash models in six weeks is a post-training loop
Look at what Google says changed. For 3.7, per Carl Franzen at VentureBeat, Google credited the three-week turnaround to “developer feedback and algorithmic improvements” and advertised “debugging and issue resolution,” “more functional web layouts and applications with fewer prompts,” and “following instructions with greater fidelity.” Andrey Kurenkov on the same release:
It is just like a refinement of 3.6 Flash rather than a new base model. It’s based on feedback from customers, supposedly, and like more RL. So you look at these benchmarks, some of them did jump a lot. Like DeepSuite, so software engineering benchmark went up from 49 to 65%. That’s a pretty big jump. Automation bench for multi-step tasks up from 17 to 30%.
— Andrey Kurenkov, Last Week in AI, 31 August 2026
For 3.8, Stevie Bonifield at The Verge:
The company claims the new model “works harder” than Gemini 3.7 Flash by performing more reasoning steps on complex tasks and “calling tools iteratively.” […] Google warns that “the model might use more tokens to maximize performance, especially at higher effort levels.”
— Stevie Bonifield, The Verge, 2 September 2026
Nobody pretrains a new base model every three weeks. A pretraining run is months of compute and data work, and its payoff is broad knowledge and raw capability. Better instruction following, iterative tool calling and harder reasoning at higher effort levels come from post-training: RL on agentic environments, preference data, reward models. The benchmarks that moved, coding and multi-step agentic tasks, are the ones RL moves. A sixteen-point jump in three weeks on the same base is a post-training result.
So the simplest explanation for three Flash models in six weeks is that Google has a base it likes and is running a fast post-training loop against it, shipping every checkpoint that clears the bar. Six months ago that was the capability everyone said Google had lost.
And the checkpoints are good. Simon Willison had 3.8 Flash build “a cool thing in HTML” in thirteen seconds for 1.8 cents, then used it to add a feature to one of his own tools; his verdict on the line is “fast, cheap, and competent at things like HTML and JavaScript.” The Register has 3.8 Flash at 59 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol and Grok 4.6, and “the cheapest model at its level of intelligence.” Bay Area Times relayed the codename, Skimaki, and the report that it “beat Anthropic’s Opus in Google’s internal coding tests.” Harris, in the same breath as his frontier-lab doubts: “this is a really good model for its tier. There’s no question.”
A Flash model level with Sol is a lab that can post-train.
What Google learned on Flash carries over to Gemini 4
Pretraining you can buy with compute and data, and Google has never lacked either. Post-training you have to learn. The old knock on Gemini, back to the 1.x days, was that the base models were good and the shipped models were worse than they should have been, hedgy and unreliable at tool use. That is a post-training complaint. The SemiAnalysis piece and the departures made everyone assume it would get worse. Over six weeks the loop instead got faster and the outputs got better on the axes that decide whether a frontier model is usable.
A small model is a cheap place to iterate on RL environments, reward models, data mixes and evals, and every improvement carries over when the big model is ready. The Flash line is the test bench. Kurenkov noted on #253 that alongside 3.6 Flash Google said it had “begun the most ambitious pre-training run yet for Gemini 4.” When that run finishes, the base gets post-trained by a team that just ran three full iterations in six weeks and knows which levers move which benchmarks.
Muhammad Ziyan at 60 Words of AI:
Anthropic ships less often and charges more. Google is going the other way. This is a bet on what wins the developer market. Speed of iteration, or peak quality.
— Muhammad Ziyan, 60 Words of AI, 2 September 2026
I don’t think those are alternatives. Fast iteration on post-training is how you reach peak quality.
Two ways I could be wrong
The Pro might be stuck for pretraining reasons. SemiAnalysis reported “industry chatter” that 3.5 Pro was “roughly Opus 4.5 level,” which would mean the base was a year behind when it came out of the oven, and post-training doesn’t close a year. Bay Area Times says Google is “months behind schedule” on a Pro-series model. If Gemini 4’s base is similarly late, a great post-training loop produces a very well-behaved second-tier model.
And small-model post-training doesn’t always transfer. RL recipes that work at Flash scale can behave differently at frontier scale, and RL compute on a frontier model is its own bottleneck. Wissner-Gross on Moonshots: “Google has other consumers fighting for their own compute internally besides AI. Whereas if you’re OpenAI or Anthropic, no, you don’t have any other non-AI users fighting for it.” Google is the only lab whose AI team argues with the ads business for chips.
Those are arguments about whether Gemini 4 will be good. Six weeks ago the argument was about whether Google could post-train at all. It isn’t any more. Kurenkov, at the end of the same episode, says: “DeepMind could just release Gemini 4 Pro and it’ll be crazy. Like 2025, they had a surprising comeback story. So we’ll see. They do have still a lot of talent. So let’s not rule them out yet.”
We’ll find out soon enough
If the gains are concentrated in agentic reliability, tool calling and coding, and it ships with effort levels that visibly spend tokens on reasoning, the Flash line was the rehearsal. If it’s context window, multimodality and benchmark tables with no behaviour story, SemiAnalysis was right and the cadence was a smaller team doing what it could.
I’d bet on the first (but not too much).