AI, ML, and networking — applied and examined.
DeepSeek V4: Flipping the $15 API Table with a Stack of Sticky Notes
DeepSeek V4: Flipping the $15 API Table with a Stack of Sticky Notes

DeepSeek V4: Flipping the $15 API Table with a Stack of Sticky Notes

Look at the bar chart in this image and you'll understand how outrageously high the blue pillar shoots up in the coding tests

It had just rained lightly in Jinan this April, and the air was filled with the cool, crisp scent of soil.

I was planning to close my laptop and go to bed early. But after a quick scroll on Twitter—oh boy, DeepSeek V4 quietly released all its APIs. No hype, not even a proper launch event.

The Hidden Threat in a Few Cents’ Bill

Putting aside buzzwords like “trillion parameters” designed to fool laymen, what actually got me out of bed in the middle of the night to run tests was its SWE-bench Verified score.

To put it simply, this benchmark throws a pile of real historical GitHub issues and tens of thousands of lines of code at you, making you understand it, find the bugs, and apply patches.

Claude Opus 4.6 previously dominated the leaderboard with a score of 80.8%. Guess what? DeepSeek V4 just surged straight to 85%.

A slight increase in accuracy might not feel that obvious. But here comes the real question.

What about the price?

Having Claude read 1 million tokens for you will cost $15. And that doesn’t even include the cost of outputting large chunks of code.

V4 costs only $0.14. Even the output is just $0.28.

$0.14—meaning the money someone else spends to run a test once is enough for you to run it a hundred times.

This isn’t some “tech-for-all” philanthropy. To be honest, this is flipping the table.

The Art of Fluorescent Sticky Notes

You might ask, how can V4 sell itself so cheaply?

Part of the credit definitely goes to the MoE architecture. For a brain with a trillion parameters, it only awakens around 32B parameters for each response. But the more crucial variable is a thing called Engram Conditional Memory.

An illustration of the Engram mechanism, which is essentially what I call "sticking post-it notes"

Let’s use an analogy. Traditional LLMs processing long texts is like having to wander through an entire massive supermarket from the first floor to the third just to find a limited-edition bag of potato chips.

Therefore, when earlier models swallowed 1 million tokens, finding details was like searching for a needle in a haystack, and extremely power-hungry.

Engram’s approach is brilliant. It’s equivalent to placing sensor-equipped fluorescent sticky notes on every aisle of the supermarket. The moment you input a request, it uses relevance signals to activate only those useful memory blocks.

The remaining massive amounts of data just lie quietly asleep in cheap DRAM memory.

To put it bluntly, everyone used to compete on who had the fastest legs. Now they’re competing on storage organization.

By the way, the underlying logic of Cursor’s recently released local codebase retrieval plugin seems to reference this kind of indexing logic too. I digress—let’s pull back.

The Stark Reality of Compute Power Cards

Putting V4 side-by-side with the current tech giants is actually quite interesting.

Besides coding capabilities, V4 has also baked vision and video generation directly into its pre-training this time. It’s not like the old days when a clumsy vision shell was forcibly strapped onto a text-based LLM.

As far as I know, by calling a single API, you can now have it write code while reading images, and even generate accompanying demo GIFs.

I just mentioned Claude 4.6, and then there’s GPT-5.3, which always seems to drip-feed its features. In terms of code generation and multimodal integration, everyone’s hard metrics are neck and neck right now.

But there’s a catch you need to know. DeepSeek’s underlying training and inference compute this time are running on Huawei’s Ascend 910C chips.

Frankly speaking, this is quite thrilling. I’ve looked up some public data, and the absolute single-card performance of the Ascend is roughly only 60% of an Nvidia H100.

That’s my superficial understanding as a layman. But they’ve relied on extreme architectural squeezing, coupled with the chip’s inherently better energy efficiency ratio, to forcibly stomp the single inference cost down to the ankles.

So, selling for $0.14 isn’t them running a loss just to gain market share.

It’s fighting this war using structural cost differences. Running a top-tier model architecture on a set of domestic compute power. Truly fierce.

If Machines No Longer Take Vacations

I’ve been pondering something these past few days.

When the cost of reading millions of lines of legacy code and fixing bugs drops to a few cents, how will us “Build-Girls” who check documentation every day write code in the future?

At this stage, writing code with AI is more like standing next to it with a whip, watching it fill in the blanks. Because you’re afraid of it hallucinating, and you’re afraid of the bill exploding after cramming the context window full.

But with a ridiculously cheap long-context monster like V4, I might just toss it 50 issues right before I get off work every day. I’ll let it spin up 50 sandbox environments and test them slowly on its own.

The next morning, I can just grab a coffee and review the results.

Think about it: what if this is used for medical record tracking or legal file organizing? Before, lawyers sifting through thousands of pages of case files was an excruciatingly physically draining task; now, for less than the cost of a latte, AI can pluck out the contradictions.

This is essentially subsidizing human attention with cheap compute power.

But the problem is, if even bug-finding and researching become zero-barrier tasks, what is left of our moat? Is it really just the ethereal soft skill of “how to ask good questions”?

It’s terrifying when you think about it deeply. Or maybe I’m just overthinking.

After all, the bizarre logic of some legacy business systems is incomprehensible even to the people who wrote it back in the day; I don’t believe AI won’t crash and burn (o_O).

It’s a blind spot in my knowledge right now. Wait until this weekend when I take my company’s ancient spaghetti code mountain for a spin and see if these so-called trillion parameters will get beaten to the ground by the wild, unorthodox methods of our predecessors.

The Rain Has Stopped

The rain outside the window sill seems to have stopped.

The potted succulent I just bought yesterday almost got completely soaked; I better hurry up and bring it inside.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *