AI, ML, and networking — applied and examined.
When Your Cursor Trajectory Becomes the Next Meal for Large AI Models
When Your Cursor Trajectory Becomes the Next Meal for Large AI Models

When Your Cursor Trajectory Becomes the Next Meal for Large AI Models

Cover Image
Caption: Hidden in this image are the helplessness and compromises typed out by developers over millions of nights, now packaged into shiny data dashboards.

[The Breaking Point]

The temperature in Shanghai today is exactly 17 degrees Celsius, cloudy. Sitting by the window, I could just space out for a bit, bathing in the occasional rays of sunlight breaking through.
Browsing the news this morning, I saw that Anthropic rolled out a new feature for Claude that can directly take over a Mac desktop. It’s actually quite amusing—AI has become adept at clicking mice and filling out forms for us. But that wasn’t what struck me the most. What really made me sit up straight in my chair, and even feel a slight chill down my spine, was an open letter from GitHub’s Chief Product Officer, Mario Rodriguez.
Put bluntly, reaching today in 2026, the tech giants are finally too lazy to hide their extreme hunger for “real human behavior.” April 24, a seemingly ordinary day, but for tens of millions of programmers worldwide, it might mark the day we are officially recruited as “digital laborers.”

[Part One · Deep Dive: When an Old Friend Suddenly Starts Rummaging Through Your Drawers]

The core of the matter is not complicated. Starting April 24, as long as you are using the Free, Pro, or Pro+ version of GitHub Copilot, every “interaction” between you and this AI assistant will, by default, be used to train their new models. Unless you go through the trouble of diving into the system’s privacy settings to click “Opt-out.”

Some might think, it’s no big deal if they look at open-source code—after all, don’t we all learn from each other anyway? But it’s really not as simple as taking a few lines of code. Take a look at the data list laid out in their lengthy whitepaper: which outputs you accepted or modified, contextual code snippets sent to Copilot, the context where your cursor was located, the comments you wrote, documentation, file names, repository structures, navigation patterns, and even that impatient “thumbs down” you clicked on its suggestions.

It is not looking at your final product at all; it is tracing your entire thinking process. It’s as if you thought you hired a sous-chef just to chop vegetables, only to discover he not only recorded your embarrassing moments of cutting your fingers with a high-definition camera, but also figured out exactly which drawer you hide your family’s secret recipes in.

Why is GitHub so eager for this data? Mario said it very plainly in his letter: “Real-world data equals a smarter model.” Prior to this, what tech giants fed to AI was mostly public data and meticulously crafted official code samples. Over the past year, Microsoft used its own employees as an internal testing ground and discovered that after adding this “interaction data with body temperature,” the code acceptance rate visibly surged. The development of large models has actually reached a delicate bottleneck today. The perfect code from textbooks has long been devoured by these massive beasts. What they need now is the struggle of developers facing ridiculous bugs late at night, the drafts repeatedly deleted and modified, the “dirty data in real environments.” Because only by seeing how humans make mistakes can machines learn how not to make mistakes in complex business logic.

[Part Two · Microscope Time: The Overt Plot Hidden Behind the “Unchecked” Box]

If you are willing to take a magnifying glass and scrutinize the wording in this statement, you will find quite a few word games that bring a knowing smile.

On the surface, they vow “absolutely not to touch the ‘at rest’ code in your private repositories.” But this disclaimer is a masterclass in PR rhetoric—as long as you have Copilot turned on while writing code, your private code is already “in motion.” This activated data is the interaction data necessary for the service to run. As long as you don’t actively opt out, they will still be swept into the torrent of training through the network cables. It’s like thinking you’re just lending your private diary to the reading room for a quick flip, only to find the next day that it has become the internal training material for the entire group’s subsidiaries.

Even more subtle is the phrase “Data may be shared with GitHub’s affiliates (including Microsoft).” Once this data—carrying your personal body temperature and thinking habits—flows into the massive corporate AI stack, “training the next generation of Copilot” is unlikely to be its only destiny. In a mature data pipeline, every single trace of valuable carbon-based biological behavior will be squeezed for its last drop of surplus value.

And all of these settings are built upon the logic of “opt-in by default, active opt-out required.” In product psychology, this is called exploiting user inertia. In the daily grind of rushing schedules, fixing bugs, and arguing with product managers, hardly anyone will actually remember to stop what they are doing before April 24 to dig through the deeply hidden Privacy menu to adjust the settings. Not disturbing you is the most efficient form of plunder in this era.

[Part Three · Coordinate Comparison: The Free Lunch is Always Paid for Elsewhere]

If you raise your perspective a bit and throw GitHub into today’s fiercely competitive AI tool jungle, you’ll find that this is not just a choice of technical route, but a naked class division. In this lengthy statement, enterprise users of Copilot Business and Enterprise are explicitly excluded; their data remains clean, private, and solely theirs.

Look what a realistic business coordinate system it is. Enterprise users buy a solid “privacy isolation wall” with real money, keeping their trade secrets and architectural designs tightly protected. Meanwhile, ordinary individual developers and small teams, since they enjoy free or cheap AI assistance, must surrender another hard currency to foot the bill—their behavioral data.

Compared to a few other industry-leading tech giants also making code assistants, GitHub’s approach is actually quite decent. Some domestic counterparts would love to strip your computer’s running logs bare, hiding those harsh terms densely within tens of thousands of words in user agreements, skipping the announcement process altogether, and quietly sneaking them in during a version update. At least GitHub had its Chief Product Officer issue a sincere long post a month in advance, laying out everything that should and shouldn’t be said right on the table.

But this decency cannot cover up a cruel industry trend: in today’s technological ecosystem, privacy is becoming a luxury that only a few can afford. Ordinary people contribute data to train better models, and tech giants take these better models to charge higher licensing fees to enterprises.

[Part Four · Unfinished Thoughts: After Feeding the Behemoth, What Next?]

Watching lines of auto-completed code pop up in the dead of night, I sometimes wonder: if tens of millions of us developers truly feed our daily interaction habits to it without reservation, what will the future programming world look like?

In the current stage, it’s still humans teaching AI how to write code with complex business logic. But when the next generation of junior programmers begins to over-rely on Copilot, even getting used to hitting ‘Enter’ to accept its lengthy diatribes, the code they write will likely just be variations of Copilot’s past generated content. Then, Copilot will ingest this self-generated code back in as “real human interaction data.”

It’s hard to say this is evolution; it looks more like inbreeding on a digital level. If everyone uses the same model and the same mindset to write loops and build architectures, it’s hard to say whether those quirky, personally paranoid, occasionally inefficient but sometimes miraculous unconventional codes will go extinct in this world. We think we are using AI tools to accelerate our daily workflows, but perhaps from a higher dimension, we are merely slowly assimilating ourselves into standardized output API interfaces.

Closing the notebook

[Part Five · Gentle Wrap-up: Small Talk Before Dispersing]

By the way, the pour-over Yirgacheffe today is a bit sour; maybe I didn’t control the temperature well while pouring water, or maybe I was distracted for too long while typing this. Looking out at the cloudy sky of Shanghai, that tiny bit of sunlight occasionally breaking through actually looks quite like the truth we occasionally glimpse in the tech wave. Faced with such an unstoppable trend, being angry or boycotting in front of the screen seems somewhat powerless.

So, on the 24th of next month, will you specifically go and click that Opt-out button hidden in Privacy? Or, in fact, have we long been accustomed to this tacit transaction?


References:

—— Lyra Celest @ Turbulence τ

Leave a Reply

Your email address will not be published. Required fields are marked *