AI, ML, and networking — applied and examined.
DeepSeek V4 Bypasses Nvidia for Huawei: The AI Hardware Shift Silicon Valley Didn’t See Coming
DeepSeek V4 Bypasses Nvidia for Huawei: The AI Hardware Shift Silicon Valley Didn’t See Coming

DeepSeek V4 Bypasses Nvidia for Huawei: The AI Hardware Shift Silicon Valley Didn’t See Coming

华为Atlas 950超节点集群与英伟达产品对比
When I first saw the data in this comparison chart, I thought my eyes were playing tricks on me.

It’s 8 degrees in Shanghai today, partly cloudy, a rather quiet Sunday. But the notifications popping up on my phone are anything but quiet—a post on Reddit updated the latest status of DeepSeek V4, saying the model could have been released last week, but was held back because of “government requirements to run on domestic hardware.” The second half of this news is the real bombshell: DeepSeek didn’t even give V4 to Nvidia for testing this time.

The First Decision That Made Nvidia Uncomfortable

Reuters’ exclusive report on February 25 stated exactly this: DeepSeek did not provide Nvidia and AMD with early access to the upcoming V4 model. Instead, it gave this opportunity to Huawei, allowing Huawei to optimize for the Ascend chips weeks in advance.

Why is this important? Because in the AI industry, providing early testing access to major chip manufacturers before a model’s release is standard operating procedure. It’s like releasing a new video game; you definitely send it to mainstream graphics card makers first to ensure compatibility. DeepSeek has also had close collaboration with Nvidia’s technical team in the past.

This time, it was cleanly cut off.

I looked it up. DeepSeek’s cumulative downloads on the open-source platform Hugging Face have surpassed 75 million. Over the past year, it drove the total downloads of Chinese open-source models to exceed the combined total of all other countries. This isn’t a small company emotionally “de-Americanizing”; this is a strategic choice made by the AI model with the highest download volume globally.

DeepSeek V4泄露的基准测试对比
V4’s leaked benchmark data—pay attention to the SWE Verified and HumanEval columns.

$6 Million vs. $100 Million: Who Do You Think Should Be More Nervous?

Let’s do the math.

DeepSeek V3’s training cost was about $6 million, utilizing around 2,000 GPUs. GPT-4’s training cost is conservatively estimated at over $100 million, using more than 10,000 GPUs. The performance gap between the two on most benchmarks is already very small.

Fast forward to the V4 generation. Leaked internal testing data shows: HumanEval (code generation capability) reached about 90%, and SWE-bench Verified (real software engineering tasks) exceeded 80%. For comparison, Claude Opus 4.6 is at 80.9%, and GPT-5.2 is at 80%. In other words, in terms of programming capabilities, V4 is already on par with some of the most expensive closed-source models in the world today.

But what about the price?

According to community estimates, V4’s API pricing might be around $0.14 per million input tokens and $0.28 per million output tokens. What is Claude Opus 4.6? $5 for input and $25 for output.

The input is 36 times cheaper, and the output is 89 times cheaper.

I previously chatted with a friend who builds AI applications. He mentioned their team’s monthly API bill was about $4,000, entirely using Claude. I did a rough calculation for him: if V4’s pricing is accurate, for the same volume of calls, the monthly fee might be under $100. His expression at that moment—how should I put it—was the “I wasted a year’s worth of money” kind of expression.

Let’s look at the timeline again. DeepSeek R1 was released on January 20, 2025. On that day, it directly triggered a $1 trillion market cap wipeout in the US tech sector, with Nvidia alone losing $600 billion. Over a year later, although Nvidia’s stock price has recovered—becoming the first company to surpass a $5 trillion market cap in October 2025—the term “DeepSeek Phobia” has already been written into Wall Street analysis reports.

If V4 truly completes its training and inference on domestic chips, the impact won’t just be about “whose AI model is stronger,” but rather, “I don’t have to use your Nvidia GPUs.”

Huawei Has Been Waiting for This Day for a Long Time

To be honest, Huawei’s Ascend chips reaching this stage hasn’t been smooth sailing.

Last August, the UK’s Financial Times reported that the DeepSeek R2 project experienced severe training delays due to insufficient multi-card interconnect speed and VRAM bandwidth of the Ascend 910C, ultimately forcing a switch back to Nvidia GPUs. During that time, there were many voices online saying domestic chips “aren’t good enough.”

But I noticed a detail: from that setback to now, in less than a year, Huawei has introduced the Atlas 950 SuperPoD.

I read through the specs of this thing multiple times. A fully loaded setup features 8,192 Ascend 950DT chips, 128 computing cabinets, and 32 interconnect cabinets, covering 1,000 square meters. Its FP8 computing power is 8E FLOPS, and interconnect bandwidth is 16.3PB/s—this figure is over 10 times the peak bandwidth of the entire global internet today.

Compared to Nvidia’s NVL144, set to be released in the second half of this year, the Atlas 950’s chip scale is 56.8 times larger, total computing power is 6.7 times, memory capacity is 15 times (1152TB), and interconnect bandwidth is 62 times. Even when compared to Nvidia’s NVL576, which is planned for market release in 2027, the Atlas 950 still leads across various metrics.

Huawei’s Rotating Chairman Eric Xu personally showcased this supernode cluster at MWC 2026, marking the Atlas 950’s first overseas appearance.

华为Atlas 950 SuperPoD海外首秀
At MWC 2026, Huawei brought this massive beast to Barcelona.

Of course, some will say that a supernode and a single chip are two different things, and Huawei’s total computing power figures achieved by stacking scale cannot be directly compared to Nvidia’s single-card performance. This point makes sense. But think about it from another angle: AI training is inherently a cluster-level engineering problem. Whoever can stably connect more computing power wins. What Huawei has done at the system integration level—developing the in-house UnifiedBus interconnect to replace NVLink, native FP8/FP4 format support, liquid cooling solutions—these are the things they truly spent immense engineering manpower and time to conquer.

A piece of information many might have overlooked: DeepSeek and Huawei Ascend have reached a “Day 0 Adaptation” collaboration model. This means that the day a new model goes online, enterprises can run it on domestic computing platforms that very day. This isn’t achieved just by signing a contract; behind it is the native refactoring of everything from operator libraries and instruction sets to the underlying architecture. According to industry test data, V4’s inference speed on the Ascend 910B has already surpassed Nvidia’s H100 by about 15%, with costs reduced by 60%.

From being “dragged down” to overtaking, in less than a year. This speed itself is an answer.

What Does This Have to Do With You and Me?

Let’s be practical.

If you are a developer, once V4 is open-sourced, you can self-host it on dual RTX 4090s or a single RTX 5090 without needing to rent expensive cloud GPUs. This means the cost for individual developers and small teams to build AI applications will drop by another order of magnitude. An independent developer with ideas can use an API for less than 100 RMB a month to access a programming model on par with Claude Opus—something unimaginable two years ago.

If you are an ordinary user, paying attention to investment opportunities in domestic AI chips might be worthwhile. The Huawei Ascend 950 is planned for launch in the fourth quarter of this year, and the entire ecosystem built around it—from chips to servers to cloud platforms—will drive demand across the domestic supply chain. While dining with a friend in the chip industry previously, he mentioned that Ascend’s ecosystem adaptation speed is several times faster than it was two years ago. Many models that could previously only run on Nvidia are now running on domestic platforms with a very similar experience.

If you are choosing AI products or API services, DeepSeek’s offerings are entirely usable right now. I personally use version V3.2 for coding and document analysis. The experience isn’t perfect, but considering the price is only a fraction of Claude’s, the feeling is like… well, never mind, just try it yourself and you’ll see.

Silicon Valley Might Not Have Realized What Happened Yet

There’s an interesting reaction from the US. A senior official in the Trump administration told Reuters that DeepSeek’s V4 might have been trained on Nvidia’s Blackwell chips, and then they attempted to “erase the technical traces of using American chips” before claiming it was trained on Huawei chips.

This accusation itself is very telling—if domestic chips really “aren’t good enough,” would you be this nervous?

CNBC recently did a 40-minute feature titled “China’s Next AI Shock is Hardware.” A year ago, they were still saying China could only play catch-up at the software layer. Now, even they admit that the hardware shock is arriving.

I finished my coffee. The sky outside the window in Shanghai is gray, but as I stare at the comparison data between the Atlas 950 and NVL144 on my screen, a thought keeps turning in my mind: the chip that was mocked for “dragging DeepSeek down” last year has evolved faster than anyone anticipated. Those who said “impossible” a year ago must be very quiet right now.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *