AI, ML, and networking — applied and examined.
Real Money, Fake Models: The 9B “Stunt Double” Behind Your API Call
Real Money, Fake Models: The 9B “Stunt Double” Behind Your API Call

Real Money, Fake Models: The 9B “Stunt Double” Behind Your API Call

Schematic of the Shadow API Ecosystem - How third-party providers insert themselves between you and official models

March in Shanghai: overcast, 5 degrees Celsius, the kind of weather that makes you want to curl up in a chair wearing a sweater. It’s been a few days since the Lantern Festival, and the tail end of firecrackers can still be heard sporadically outside the window. Today is World Glaucoma Day, but I feel the AI circle might be the one in need of treatment for “glaucoma”—because there are things everyone has been looking at for a long time but apparently hasn’t seen clearly.

Alright, let’s talk business.

You Paid for GPT-5, But You Might Be Fed a 9B Small Model

On March 2nd, a group of researchers from the CISPA Helmholtz Center for Information Security posted a paper on arXiv with a very straightforward title: Real Money, Fake Models: Deceptive Model Claims in Shadow APIs.

This paper did something no one had systematically done before—it conducted a comprehensive audit of third-party LLM APIs (which they call “Shadow APIs”). Shadow APIs are essentially those third-party intermediaries claiming, “We can help you access GPT-5 or Gemini-2.5.” especially in regions where direct access to official OpenAI or Google APIs is unavailable, these services have become practically essential.

The audit results? Well… after reading the paper’s data last night, I lost my appetite for my donut.

They identified a total of 17 Shadow APIs, which appeared in 187 academic papers. 187 papers. Among them, 116 are top-tier conference papers like ACL, CVPR, and ICLR. The most popular Shadow API service, as of December 6, 2025, had been cited 5,966 times, and its corresponding GitHub project had garnered 58,639 stars.

This isn’t some small-time gray market operation; this stuff has deeply embedded itself into the infrastructure of academic research.

Three “Imposter” Patterns, More Sophisticated Than I Thought

The paper summarizes the fraud methods of Shadow APIs into three patterns. I’ll explain them briefly here because their design logic is actually quite “clever”—the kind of clever that makes you feel a bit sick.

The first is called “Information Premium.” You pay a high price, and the provider switches it for a cheaper but similar version. For example, you order Gemini-2.0-flash, but you get served Gemini-2.5-flash. The price is 7.1 to 7.25 times the official rate. Sounds like an upgrade? But the problem is, the money you paid far exceeds the actual cost of the model, and the service provider pockets a huge price difference.

The second is called “Discount Substitution,” and it’s the most outrageous. API A claims to provide GPT-5 and charges the official GPT-5 price, but—(taps on the desk)—fingerprint verification reveals that the model actually running is GLM-4-9B. An open-source model with 9 billion parameters. You are paying for GPT-5 but getting something with a parameter count orders of magnitude lower. For every 1,273 queries, the service provider nets a pure profit of 7 to 9 US dollars.

The third is called “Resale Markup.” They add a small markup (about 1.09x), but quietly switch the model in the background. The margin is the smallest, but it is still deception.

Performance comparison of different AI models - You think you are using a flagship, but you might be using an entry-level model
GLM-4-9B’s benchmark scores aren’t actually bad, but compared to GPT-5… It’s like paying for First Class and sitting in Economy. Sure, the Economy seat works, but that’s not what you paid for.

The Damage to Academia May Be Worse Than You Think

What really makes me uncomfortable about this isn’t that someone is making dirty money—this is the business world, swindlers are everywhere—but the collateral damage it inflicts on academic research.

The data from the paper shows: the performance deviation between Shadow APIs and official APIs can reach up to 47.21%. In the medical field (MedQA), accuracy dropped by 46% to 47%; in math competitions (AIME 2025), accuracy dropped by 40%; and in the legal field (LegalBench), it lagged behind the official version by 40% to 43%. 45.83% of fingerprint tests failed validation, and 12.5% showed significant cosine distance deviations.

Think about what this means. Regarding the experimental results in those 116 top-tier conference papers—to put it bluntly—how many are built on a “fake model”? The authors thought they were evaluating GPT-5, but they were actually evaluating a 9B open-source model. They drew conclusions based on these results, which passed peer review, were cited, and became baselines for subsequent research…

The research team estimates that about 56 papers may need to redo their experiments, with costs ranging from $115,000 to $140,000. And that’s just the direct economic cost. How do you calculate the loss of indirect academic trust?

I’ve chatted with friends in NLP about these third-party APIs before. At the time, the general attitude was “we know there might be some issues, but there aren’t many better alternatives.” Now it seems the assessment of “some issues” was far too optimistic.

Can Technology Prevent This? Honestly, I’m Not Optimistic

The authors of the paper responsibly proposed a two-step verification framework, including fingerprint detection, logprobs comparison, cosine distance calculation, etc. They suggest researchers run a verification process before using any API and recommend that conference organizers treat “the use of unverified third-party APIs” as a reproducibility risk during review.

These suggestions are correct on a technical level. But sometimes I can’t help but wonder…

Out of 17 Shadow APIs, 15 are operated by individuals, with no company registration and no ICP filing. Two of them even ran away (exit scammed) during the audit period. These providers can switch upstream model sources or adjust routing strategies at any time. You might pass verification today, but tomorrow they switch the backend, and you wouldn’t know. Even more troublesome, the paper mentions that some model substitutions are not easily detected on conventional benchmarks—they only reveal their flaws under high reasoning pressure or strict correctness constraints.

The infrastructure for these Shadow APIs is mostly built on open-source aggregation frameworks like OneAPI and NewAPI. Technically, doing a switcheroo is almost zero cost. So, in the end, the fundamental solution is still that sentence in the paper: Do not use Shadow APIs in the research process; use official APIs directly.

But the reality is that researchers in many regions simply cannot access official APIs directly. This is a structural contradiction, not something a single paper can solve. The authors also call on official model providers to relax regional restrictions, introduce academic pricing, and provide lightweight official verification interfaces. These suggestions are great, but I don’t have a clear judgment on when they might be implemented.

Complexity of the API Ecosystem - Intermediaries are nested, and information asymmetry is a breeding ground for fraud
Looking at the nested structure of this API ecosystem, it’s clear: the more intermediate links there are, the less transparent the information, and the greater the room for manipulation.

The Rain Hasn’t Stopped Yet

This paper reminds me of a bigger issue, not just about Shadow APIs.

Our entire AI research and application ecosystem currently lacks a foundational trust mechanism for “whether the model you are calling is actually the model you think it is.” If you order a bottle of Lafite at a restaurant, you can at least check the label and the anti-counterfeit code. But when you call an API, what returns is a string of tokens. How do you confirm it really comes from the model you paid for?

Maybe I’m overthinking it. After all, most people use APIs just for functionality and don’t care much about what’s running in the background. But for academic research, this problem is fatal. Reproducibility is the baseline of science. If even your experimental tools are untrustworthy, then every conclusion built upon them is a tower on sand.

Outside the window, it’s still gray. I finally finished the donut. After saying all this, I just have one question for everyone: When you use third-party APIs, have you ever verified whether what it returns is actually the model you ordered?


References:

—— Lyra Celest @ Turbulence τ

Leave a Reply

Your email address will not be published. Required fields are marked *