AI, ML, and networking — applied and examined.
Training AI on Your Code with Zero Refunds: The Hidden Trap in Cheap Coding Plans
Training AI on Your Code with Zero Refunds: The Hidden Trap in Cheap Coding Plans

Training AI on Your Code with Zero Refunds: The Hidden Trap in Cheap Coding Plans

The hidden trap in the purchase interface of a domestic Coding Plan
(This illustration is quite vivid; it’s indeed easy to get dizzy looking at these dazzling subscription plans nowadays.)

It’s raining outside again. That’s just how March in Shanghai is, making one’s mood a bit damp as well. While brewing coffee this morning, I routinely scrolled through developer communities, only to be jolted fully awake by the Terms of Service of a “Coding Plan.”

The Devil is in Clause 5.2.2

The origin of this story is that several domestic tech giants have recently, as if by prior agreement, launched “Coding Plans” (code assistance subscriptions) targeted at programmers. The prices are incredibly tempting—some cost just a few cents for the first month, others about forty or fifty RMB a month—granting you unlimited access to the top-tier domestic Large Language Models directly within your IDE.

However, in that “Special Agreement” you usually check without reading, there lie two glaring paragraphs. I’ll quote them word for word for you:

5.2.1 You acknowledge and agree that once the Coding Plan service is purchased, unsubscription and refunds are not supported.
5.2.2 You agree and authorize us and our affiliates to store and use the content you input by calling the model and the content generated by the model during your use of the Coding Plan (“Coding Plan Data”) for service improvement and model optimization. If you wish to terminate the authorization of the Coding Plan Data, you may do so by ceasing the use of the Coding Plan service, but the scope of termination does not cover the Coding Plan Data you have already authorized us to use.

Translation: As long as you buy it, you can’t get a refund; as long as you use it, your prompts, the error codes you paste in, and the code the model generates for you—we’re taking it all to train our models. If you ever regret it and want to stop, sorry, we’ll keep using the data we’ve already taken. No take-backs.

This is fascinating. You see, in the privacy statements for their general API services, these giants usually state in bold red letters, “We will never use your data for model training.” But when it comes to the cheap and convenient Coding Plans, the tone abruptly changes. Not only will they use it, but their stance is also extremely uncompromising.

Who is Fleecing Whom?

Let’s do the math.

Why are these tech giants willing to bundle top-tier models—which usually cost dozens or hundreds of RMB when billed by tokens—and sell them to you for a “cabbage price” monthly flat rate? Are they really doing it to democratize technology?

To put it bluntly, it’s like going to a salon for a heavily discounted hair wash, but the catch is the barber gets to test a new pair of clippers on your head. The training of Code LLMs has hit a bottleneck. All the open-source public code on GitHub that could be scraped has long been scraped clean by everyone. What major tech companies lack the most right now is high-quality “real-world business scenario data” and “highly logical, complex human prompts.”

And you, writing code in your IDE, happen to be the perfect machine to produce this golden data.

In order to fix a bug buried deep in your business logic, you converse back and forth with the model, providing context with real variable names and database table structures. Finally, the model provides an elegant solution, and you click “Accept.” For Reinforcement Learning training, this complete interaction pipeline is an incredibly precious, top-tier corpus.

Pricing and model list of a tech giant's Coding Plan
(Just look at this pricing table; being able to access so many top-tier models for a few dozen RMB means the underlying cost calculation is definitely playing a different game.)

You think you spent a little money to buy a month’s access to an advanced tool. In reality, the tech giant only gave up a tiny server fee to hire you, a senior software engineer, to do a month of free data annotation for them. What’s more, backed by the “no unsubscription and no refund” Clause 5.2.1, regardless of whether you end up hating the experience halfway through, it’s a guaranteed win for them.

Same Code Assistants, Where’s the Gap?

If you look horizontally across the industry, the perception of this issue gets even more complicated.

Let’s see how Cursor, which currently dominates the market, handles this. For Free or Pro users, its privacy mode is also turned off by default, and it indeed collects code for optimization. But for Enterprise (Business plan) users, they implement a ZDR (Zero Data Retention) policy, and have signed strict non-disclosure agreements with backend providers like OpenAI and Anthropic, explicitly guaranteeing that user code will not be used to train models.

In contrast, the current domestic Coding Plans, though cheaper by orders of magnitude, offer almost a one-size-fits-all “take it or leave it” approach to privacy options.

Moreover, that “no unsubscription and no refunds” clause has caused trouble in actual execution. Last month, when a leading domestic vendor upgraded its code model, the switch between the old and new models led to a wave of degraded user experiences. At that time, a large number of users who had prepaid for annual subscriptions demanded refunds, only to slam right into the brick wall of this “purchase constitutes confirmation” clause. Finally, under immense public pressure, the company reluctantly opened a special compensation channel to handle it.

To be blunt, companies wanting to harvest data when rolling out cheap services is nothing new in the internet age. But you must understand the problem here: completely cutting off all fallback options for users with such draconian legal terms, not even granting the right to “revoke historical data authorization,” is somewhat arrogant.

API errors and limits in FAQs
(By the way, take a look at these backend error interception prompts. Often, your requests have already been parsed through multiple layers before they even reach the model.)

Is Your Code Really Still Yours?

Last weekend, I grabbed tea with a friend who works in enterprise security at a major financial institution, and we chatted about this.

Model configuration of IDE accessing Coding Plan
(When configuring the API keys for these big tech models in your IDE, you are essentially opening an outward-facing window for your codebase.)

I sometimes wonder: if you are an independent developer or a programmer in a small startup, and you use this kind of Coding Plan to save money, pasting the company’s core recommendation algorithm or encryption logic into the chat box… After this piece of code is absorbed by the cloud, will there come a day when your competitor asks the model, “How do I implement a similar recommendation feature?” and the model just spits out your source code verbatim?

Maybe I’m overthinking it. The security teams at these major tech companies will surely say: we do strict data anonymization before ingestion, we scrub out hardcoded passwords, Tokens, and specific naming conventions, and we avoid privacy leaks by fragmenting and restructuring the data.

But can it technically be made absolutely clean? I’m not so sure. The “memorization effect” of Large Language Models has always been a major unresolved challenge in academia. Sometimes the model just rote-memorizes a highly distinctive snippet of text and gets triggered by specific prompts.

The day before yesterday, while running a local data cleaning script, I suddenly realized an awkward dilemma: to improve the accuracy of the AI-generated code, developers must provide it with enough business context; but once you provide the context, you are essentially handing over the intellectual property of your code to the cloud. For those teams that don’t yet have the budget to purchase expensive enterprise-grade private deployments, this is almost an unsolvable dead end.

Honestly, looking closely at this agreement, reading between the lines reveals the tech giants’ extreme thirst for high-quality data and their tentative steps towards commercial monetization. I’m not against trading data for convenience—after all, there is no free lunch in this world, and cheap lunches usually have their hidden prices clearly marked.

(Sighs) The coffee has almost gone cold. I’m going to reheat it. When you all write code, remember to mask your sensitive data.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *