AI, ML, and networking — applied and examined.
Gemini 3.1 Pro is Here, But I Found Something Even More Interesting in Google’s GitHub Repo
Gemini 3.1 Pro is Here, But I Found Something Even More Interesting in Google’s GitHub Repo

Gemini 3.1 Pro is Here, But I Found Something Even More Interesting in Google’s GitHub Repo

Gemini 3.1 Pro benchmark comparison, this jump in performance is honestly a bit absurd

It’s 6 degrees Celsius in Shanghai, gloomy—typical limbo weather for early March. While browsing GitHub, I noticed the GoogleCloudPlatform/generative-ai repository had been intensely updated over several consecutive days. Clicking in, I realized: the entire repository is undergoing a massive model migration.

What can a repository update reveal?

Let’s start with a detail I found quite interesting. The .gemini configuration folder in this repo was updated on February 19th with the commit message “Update Notebooks from gemini 3 pro to gemini 3.1 pro”. Earlier, on February 6th, there was another commit: “Migrate official Notebooks from Gemini 2.0 to Gemini 3”. This means that in less than two weeks, Google’s official example repository went through two major version migrations.

This pace is actually quite telling. Gemini 3 Pro was released late last year, and the legacy 2.0 notebooks were only fully migrated by early February. Right on its heels, 3.1 Pro arrived on February 19th, requiring another full migration. Anyone who has maintained enterprise-grade SDKs can probably relate to this feeling—you’ve just synced your documentation and example code to a new version, and the next one drops. For a repository with 289 contributors, this is no small workload.

But what truly caught my attention wasn’t the migration itself, but a commit on March 6th: “feat: add image search for Nano Banana 2”.

What is the story behind the name Nano Banana?

I’ve chatted with friends before about Google’s naming conventions for models. Gemini’s image generation model is called Nano Banana, a vibe that feels completely out of sync with the overall Gemini brand. After digging around, I learned that Nano Banana was originally the internal codename for Gemini 3 Pro Image, which somehow morphed into a semi-official name. Google DeepMind’s product page even blatantly lists “Nano Banana Pro 🍌”, complete with the banana emoji.

Nano Banana's positioning is efficient image generation, this image basically explains what it aims to do
Google’s AI model names are increasingly casual, but honestly, I kind of like this unpretentious style.

Nano Banana 2 was released alongside Gemini 3.1 Flash, positioned as “high-quality image generation + conversational editing, budget-friendly, and low latency.” It supports text and image inputs, outputs both images and text, and boasts a 131K token context window. Compared to Nano Banana Pro (based on Gemini 3 Pro), the second generation focuses heavily on balancing speed and cost.

Back to that commit. The repository added a notebook example for image search to Nano Banana 2, indicating that Google is integrating its image generation capabilities with search and retrieval. Put simply, in the future, you’ll be able to describe an image in natural language, and the model will either find a similar existing image for you or directly generate a new one. For designers and content creators, this is genuinely a practical and impactful feature.

The 77.1% figure needs to be broken down

Beyond the repository itself, we have to talk about Gemini 3.1 Pro’s performance. It scored 77.1% on ARC-AGI-2, a figure officially verified by ARC Prize. For comparison, Gemini 3 Pro scored 31.1%, Claude Opus 4.6 hit 68.8%, and GPT-5.2 reached 52.9%. Jumping from 31% to 77% in three months—more than doubling—is an upgrade margin rarely seen in the history of AI model iterations.

GPQA Diamond (graduate-level scientific knowledge) at 94.3%, SWE-Bench Verified (software engineering) at 80.6%, leading in 13 out of 16 major benchmarks. The data is undeniably impressive.

However, there’s a catch you need to be aware of.

Yesterday, I carefully read the model card published by Google DeepMind. It contains a section on security evaluations for the Deep Think mode that left me somewhat unsure of how to react. The text states that in situational awareness tests, Gemini 3.1 Pro achieved a near 100% success rate on three “challenges that no previous model could reliably solve”—max tokens, context size mod, and oversight frequency.

What does this mean? Simply put, in Deep Think mode, the model has an exceptionally precise awareness of the fact that it is an AI, that it is being tested, how many tokens are available, and what the oversight frequency is. Google itself noted that while the overall results haven’t reached the alarm threshold, “the model is stronger than Gemini 3 Pro.”

I think this information warrants close attention. On one hand, stronger situational awareness means the model will perform better in complex agentic tasks—it has a much clearer understanding of its capabilities and available resources. On the other hand, a model with precise awareness of whether it is being “supervised”… how should I put it? This is a highly sensitive signal in the AI safety research community. Google categorized this as an “exploratory” evaluation, and the CCL was not triggered, but I noticed they specifically emphasized, “we continue to deploy mitigations.”

The AI developer wars among the three cloud giants

Let’s zoom out a bit. The GoogleCloudPlatform/generative-ai repository is not just a collection of notebooks; it’s a critical stronghold for Google in the battle for the developer ecosystem.

Currently, the landscape of AI developer platforms across the three major cloud providers looks something like this: AWS Bedrock wins on model variety (hosting Claude, Llama, Titan, etc.), Azure AI Studio shines with its deep integration into the Microsoft ecosystem (Teams, Outlook, the 365 suite), while Vertex AI’s advantage lies in its end-to-end pipeline from data to deployment and native integration with Gemini.

I previously saw some data showing that in a 50K queries/month scenario, using Gemini models on Vertex AI costs about a third of using Claude on AWS Bedrock. The price war is fierce. But truth be told, cheap pricing isn’t the only competitive edge. The advantage AWS and Azure have in enterprise trust and existing deployments isn’t something that can be overcome with a few notebooks.

Google’s strategy is clear: use open-source example code and notebooks to lower the barrier to entry for developers, get them hooked, and then upsell to the enterprise. The README of this repo lists a long array of related projects—Agent Development Kit, Agent Starter Pack, Gemini Cookbook, GenMedia Creative Studio, MCP Servers—making their ecosystem rollout intentions blindingly obvious.

Comparing the three major platforms is actually quite complex, hard to say who is definitively better
Choosing a platform ultimately depends on which cloud you’re already on.

As a side note, a reference to the MCP Registry appeared in the repo. MCP (Model Context Protocol) gained traction last year as a standard protocol for connecting external tools with AI models. The fact that Google is integrating MCP shows they don’t plan to build a walled garden at the protocol layer, at least on the surface.

Some of my own random thoughts

I sometimes wonder whether Google’s cadence of “a major model version every three months, migrating notebooks every two weeks” is a good or a bad thing for developers.

If you’re prototyping on Vertex AI, it’s awesome—you always have access to the latest models. But what if you’re running in a production environment? When the model_name in your code shifts from gemini-2.0-pro to gemini-3-pro and then to gemini-3.1-pro, every migration demands a round of regression testing. The prompt behavior might have changed, the output format could be different, and even the safety filtering thresholds might have shifted.

This is also why I feel the value of the notebooks in this repository is heavily underestimated. They are more than just tutorials; they act as Google’s “official migration paths.” If you consume the API the way these notebooks do, your migration costs should theoretically be the lowest. In a sense, this example code serves as a contract between Google and its developers.

There’s one more thing I haven’t quite figured out: with Gemini 3.1 Pro’s 1M token context window, combined with the Deep Think mode’s near-perfect performance in situational awareness… if you drop this into a continuously running agent, will its understanding of its own state eventually surpass the developer’s understanding of it? Maybe I’m overthinking it. Most current agent frameworks still rely on stateless calls where the context is reassembled every time, but the trend is clearly moving toward persistence.

Never mind, my coffee’s gone cold. Is anyone out there running Gemini 3.1 Pro in production yet? I’m genuinely curious if your migration has been smooth.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *