AI, ML, and networking — applied and examined.
Google’s AI Iteration Anxiety: What a GitHub Repo’s Commit History Reveals about Gemini 3.1 Pro
Google’s AI Iteration Anxiety: What a GitHub Repo’s Commit History Reveals about Gemini 3.1 Pro

Google’s AI Iteration Anxiety: What a GitHub Repo’s Commit History Reveals about Gemini 3.1 Pro

Gemini 3.1 Pro release timeline; the iteration pace from Gemini 3 Pro to 3.1 Pro is indeed very fast
This timeline looks quite neat, but developers caught in the middle probably don’t feel that way.

March in Shanghai, 6 degrees, overcast, that kind of grayish sky outside the window. While scrolling through GitHub, I noticed a repository’s commit history that was quite interesting—specifically, GoogleCloudPlatform/generative-ai.

Three Months, Two Major Migrations

You might not be unfamiliar with this repository—Google Cloud’s official generative AI sample code base, filled with notebooks, demos, and sample apps, backed by 289 contributors. Yesterday, while flipping through its commit history, I noticed a detail: on February 6th, there was a commit titled “Migrate official Notebooks from Gemini 2.0 to Gemini 3”. Then, just 13 days later, on February 19th, came another: “Update Notebooks from gemini 3 pro to gemini 3.1 pro”.

13 days.

This means Google’s own official sample code, barely migrated from 2.0 to 3.0 before the ink was even dry, had to be revised all over again. If you are a developer who relies on these notebooks to learn and build projects, you can probably imagine the feeling—the code you just got running last week might need an API change this week.

Frankly, this isn’t just about version numbers. Google broke its previous convention of using .5 for mid-term updates and, for the first time, released a .1 incremental update. The official narrative is a “focused intelligence upgrade”, aiming specifically at boosting reasoning capabilities. The numbers do look impressive: 77.1% on ARC-AGI-2, more than doubling the score of 3 Pro; APEX-Agents jumped from 18.4% to 33.5%; SWE-Bench Verified hit 80.6%, basically tying with Claude Opus 4.6’s 80.8%.

Comparison data of various models across multiple benchmarks
The data density in this comparison table is high; pay attention to the ARC-AGI-2 row, where the gap is the most obvious.

What Really Bothers Me Isn’t the Benchmarks

JetBrains’ AI Director Vladislav Tankov said something quite pragmatic: “We observed up to a 15% improvement over the best performance of Gemini 3 Pro Preview. The model is stronger, faster, and more efficient, requiring fewer output tokens to yield more reliable results.” The phrase “fewer output tokens” is the key here—for API users, this directly translates to lower costs.

But if you open Google’s developer forums, the vibe is quite different. One thread is directly titled “Gemini Pro 3.1 is kind of garbage for coding in AI Studio”. Although the title is a bit emotional, the discussion underneath isn’t purely complaining. Another thread is titled “Gemini 3.1 pro was good a week ago, but now it’s completely down”, posted on February 27th. An even earlier one, “Gemini 3 review after 1 month: inconsistent at best, poor at worst”, has over 1600 views.

The gap between benchmarks and actual experience is nothing new in the LLM space. But there is something unique happening on Google’s side: they are doing two major things simultaneously—pushing the new Gemini 3.1 Pro model and forcing an SDK migration. The Vertex AI SDK will be deprecated this June, and everyone must migrate to the Google Gen AI SDK. Someone in the LangChain4j community opened an issue discussing this back in January, and by late February, developers were still reporting that the Gemini 3 series models couldn’t properly use the thinking budget parameter in the old SDK at all.

To put it bluntly: Google is asking developers to chase both model versions and SDK versions at the same time, running on two tracks simultaneously.

What Are the Neighbors Doing?

Looking laterally. Claude Opus 4.6 scored 80.8% on SWE-Bench Verified, virtually indistinguishable from Gemini 3.1 Pro’s 80.6%. GPT-5.2 sits at 80.0%. In the dimension of actual code bug fixing, these three giants have essentially entered a stalemate.

However, there are still differentiators. Gemini 3.1 Pro’s default context window is 1 million tokens, roughly 2.5 times GPT-5.2’s 400k, and 5 times Claude’s default 200k. In terms of multimodal inputs, Gemini is also the most comprehensive, supporting text, images, code, audio, video, and PDFs. Moreover, Gemini 3.1 Pro is priced significantly lower than Claude Opus 4.6—this is crucial for enterprise users running high volumes.

But there’s an issue you should be aware of: Gemini’s knowledge cutoff date is January 2025, whereas GPT-5.2’s is August 2025, and Claude’s is also around mid-2025. An information gap of over half a year will affect output quality in certain scenarios.

There is also Google’s Antigravity platform—this is the “agentic development platform” they launched last November, akin to a VSCode fork that can spin up multiple AI agents to collaborate autonomously on development tasks. The concept is incredibly cool, but the feedback I saw earlier was: paid users encountering multi-day quota lockouts, frequent crashes, and uncontrollable agent behaviors. It was still in public preview status as of January this year. I chatted with a friend about it; he said he tried it once and went back to using Cursor.

Three-sided architectural design of Google Antigravity
The transition from traditional AI-assisted coding to autonomous agent coding is a good idea, but stability is another matter.

Some Things I Wonder About Sometimes

Back to that GitHub repository. 289 contributors, constantly migrating notebooks from one version to another. On March 6th—just yesterday—the latest commit was “feat: add image search for Nano Banana 2”. Nano Banana is Google’s image generation model, used for image editing within Antigravity. It’s a pretty cute name.

But what I’m pondering is a bigger question: Google’s AI product line now spans Gemini (consumer and API), Vertex AI (enterprise), Antigravity (developer platform), Google AI Studio, NotebookLM, Android Studio integrations… The model versions on every single line are iterating rapidly, and the SDKs are changing too. For a team looking to seriously build products within the Google ecosystem, which track should you bet on?

If the Vertex AI SDK really sunsets in June, and the Gen AI SDK isn’t mature enough yet—will there be issues during this transitional window? Perhaps I’m overthinking it. With Google’s massive enterprise customer base, if something really goes wrong, they will surely extend the deadline. But that underlying anxiety of “the code might need to change at any moment” is a sentiment repeatedly echoed by many developers in the community.

By the way, I noticed the .gemini folder in that repository—yes, Google’s own repository has its own AI configuration files. In a sense, this is AI helping humans write code about AI, and then this code teaches other humans how to use AI. Thinking about this makes me feel… how should I put it.

My coffee has gone cold. Let’s wrap it up here for today.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *