DeepSeek just gave away the software layer Nvidia's China moat stands on
Six open-source Ascend modules, including a CUDA rival called TileLang, move China's chip software bid from roadmap to repository. For Nvidia, the reopened-China bull case now rests on Beijing's permission alone.
Richard Tang · 3 min read
DeepSeek open-sourced six software modules for Huawei's Ascend AI chips on Wednesday 1. The package goes at the one layer Beijing's chip drive could not simply buy: the software that makes a chip programmable. The Huawei production-line wager already reported here has now produced shippable code.
Bloomberg calls the TileLang release China's answer to CUDA, free to download like everything else in it 2. As of Wednesday, the TileLang repository officially lists Huawei's Ascend 950 as a supported backend, with native code generation and automatic scheduling 3.
The kernels were the bottleneck, not the transistors
AI chips are only as good as their kernels, the hand-tuned matrix-multiply and attention code where performance is won or lost, and that craft has lived inside CUDA for a decade. TileLang collapses the work: write a kernel once in Python-like syntax, and the compiler generates optimized code for Nvidia, AMD, Apple silicon and now Ascend 3.
DeepSeek paired it with DeepGEMM for matrix multiplication and DeepEP for shuttling data between the 128 Ascend 950 processors of the Huawei supernode it tuned on, reporting performance that approached the hardware's limits in its own tests 14.
DeepSeek's own workloads already run on it
This is not a demo stack. TileLang examples for DeepSeek's V4 workloads have been in the repository since May 3, the lab has run V4 inference on Ascend since April with Huawei's full support 1, and Huawei supplied extensive technical help on the software 4.
Huawei says Ascend 950 supernodes are already in commercial use and that more than 1,000 older 910C supernodes are deployed 5, while the 960DT training chip has been pulled forward to the first quarter of 2027, three quarters early 1. Inference today, training from next year: the stack now covers both halves of a Chinese lab's compute bill.
Domestic silicon was selling before the software existed
Demand was never the constraint. Cambricon's first-half revenue rose 108% to 6 billion yuan, while Hygon forecast up to 9.3 billion yuan and Moore Threads up to 1.75 billion for the same six months, carried by Beijing's self-sufficiency push and US export controls 6. Cambricon, Moore Threads and Alibaba's T-Head had already confirmed their chips run DeepSeek's models 1.
The telling part is that the revenue arrived before the tooling did: buyers were not waiting on software to place orders. TileLang's repository already carries backends for Moore Threads and Hygon hardware 3, and Wednesday's release supplies the kernels, compiler and communication libraries for Ascend 1. For a Chinese buyer choosing between metered Nvidia hardware and domestic silicon, the last reason to wait is gone.
Domestic chip vendors were scaling before DeepSeek's software arrived
- H1 2026 revenue
- Estimate
Data
| H1 2026 revenue | |
|---|---|
| Hygon (estimate) | ¥9.3B |
| Cambricon | ¥6B |
| Moore Threads (estimate) | ¥1.75B |
Nvidia's China case now rests on permission alone
Nvidia guided $108 billion for the current quarter assuming no data center compute revenue from China 7. Jensen Huang has said Nvidia's China AI chip share fell from about 95% to zero 7. It still billed $7.9 billion to China-headquartered customers last quarter, almost none of it AI compute 7.
The bull case for a reopened China has two legs: Beijing's meter, reportedly being tested with possible RTX Pro 5500 approvals for ByteDance and Alibaba 8, and CUDA's supposed unclimbability. The first is politics. The second now has a multi-vendor counterexample in a repository. CUDA's tooling still wins outside China, but the software gap was the one bear case export controls could not close, and this week it closed in code.
Watch for two confirmations: a Chinese lab putting a frontier training run on Ascend, and the 960DT arriving in the first quarter as promised 1. One is a lab's choice; the other is a date.
The reward goes to whoever made the hardware programmable
DeepSeek and Huawei did the middle work themselves, in the open, and they are the ones who should capture the value: not whoever sits closest to the transistors, but whoever made the hardware programmable. The moat was always software. China's answer is now public, and free.
More about NVIDIA
NVDA · Fiscal Q2 2027Revenue rose 106% to $96.2B, 92.5% of it from Data Center, “driven by the ramp of our Blackwell Ultra infrastructure”. Gains on equity stakes of $7.8B lifted net income to $59.7B.
Show as a table
| Line | Value |
|---|---|
| Revenue | $96.2B |
| Gross margin | 75.0% |
| Operating margin | 66.2% |
| Net income | $59.7B |
| Hyperscale | $48.7B |
| AI clouds & enterprise | $40.3B |
| Data Center | $89.0B |
| Edge Computing | $7.2B |
| Gross profit | $72.1B |
| Cost of revenue | $24.1B |
| Other income | $7.8B |
| Operating income | $63.7B |
| Operating expenses | $8.4B |
| Tax | $11.8B |
| R&D | $7.1B |
| SG&A | $1.4B |
Deepdive
AI-generated from this story and its cited sources. Not investment advice.



