On September 30, in a post on its official WeChat channel quickly picked up by Reuters and outlets worldwide, the Chinese AI lab DeepSeek did something with no precedent in the AI chip wars. It gave away the good stuff. Working with Huawei, the company open-sourced a six-module software toolkit built for Huawei's Ascend 950 AI accelerators, with a high-level kernel language called TileLang at its center.

Chip companies hoard software. Nvidia built the most durable competitive position in modern technology less with its GPUs than with CUDA, the platform that has kept millions of developers building exclusively on Nvidia silicon for nearly two decades. DeepSeek's move is a direct assault on that moat, and it matters less for what it proves today than for what it starts.

The release: six modules, one idea

The toolkit mirrors the open-source tools DeepSeek had previously released for Nvidia hardware. TileLang is the centerpiece: a domain-specific language for writing high-performance AI kernels, the tight inner loops of computation that decide how fast a model trains or answers. Around it sit five supporting modules: DeepGEMM-Ascend for matrix multiplication, DeepEP-Ascend for communication across devices in a cluster, TileKernels for vector operations and memory access, FlashMLA for long-context processing, and DeepSelect for data filtering.

The design choice that matters is parity. DeepGEMM-Ascend keeps the API and workflow of DeepGEMM on Nvidia hardware, and TileKernels exposes the same Python APIs across Nvidia GPUs and Huawei NPUs. The goal: move a workload between the two by changing a configuration flag rather than rewriting kernel code. The companies also jointly tuned a 128-chip Ascend 950 supernode for computation and inter-chip communication.

TileLang itself predates this moment: it was developed by Peking University researchers, published as an academic paper in April 2025, and has been in use inside DeepSeek for roughly a year. As of September 30 it officially supports the Ascend 950 as a backend, joining Nvidia's CUDA, AMD's ROCm, and Apple's Metal.

Why the moat is software, not silicon

Enjoying this story?

Get the five most important stories in tech, every morning. Free.

To understand what is being attacked, you have to understand what Nvidia actually owns. At its GTC 2026 conference, Nvidia said more than six million developers now build on CUDA, up from 4.5 million a year earlier, supported by a library of more than 400 GPU-accelerated tools including cuDNN, cuBLAS, and NCCL. CUDA launched in 2007: nineteen years of libraries, debugging tools, forum answers, and institutional knowledge, all running only on Nvidia hardware.

Moving an existing codebase off CUDA has historically meant re-implementing kernels, debugging unfamiliar compilers, and accepting performance regressions wherever Nvidia's libraries are hand-tuned. Export controls aimed at starving China of AI chips were supposed to widen that gap. Instead, they may have focused Beijing's attention on the one part of the stack where a substitute is actually constructible: the software layer above the silicon.

Nvidia seems to know it: its most recent annual filing stated that competitors have built larger developer and customer ecosystems to challenge it worldwide. In a May 2026 interview, CEO Jensen Huang acknowledged that US export restrictions had cost the company roughly $50 billion in China this year alone. "We have largely conceded the China market to Huawei," he said. The hardware fight in China, for now, is effectively over. The software fight is just beginning.

CUDA's lock-in was never really about silicon. It was about nineteen years of code that nobody wanted to rewrite, and that is exactly what DeepSeek is trying to make irrelevant.

What TileLang actually does differently

The CUDA Moat, by the Numbers

What TileLang is up against, and what DeepSeek brings to the fight.

CUDA developers
6M+
CUDA accelerated libraries
400+
CUDA ecosystem age
19 yrs
TileLang backends
4
Open-sourced Ascend modules
6
Supernode test chips
128

Note: For illustrative purposes only.

CUDA asks developers to think in threads: individual GPU threads, manual memory access, manual synchronization, all of it coupled to Nvidia's specific architecture. It gives experts maximum control and maximum portability pain. TileLang operates one level up. The fundamental unit is the tile, a block of data the compiler reasons about as a first-class object. Developers describe what computation to perform on which data, and the compiler handles scheduling and synchronization on the target hardware. The same TileLang kernel can then emit CUDA for Nvidia, Ascend-C code for Huawei, HIP for AMD, or Metal shaders for Apple.

This is the trade every abstraction offers: less fine-grained control in exchange for portability and a smaller development burden. DeepSeek calls TileLang its "core tool" for AGI research. The claim that it offers "a simpler programming model" than CUDA while reaching the hardware's full performance potential is ambitious, but it is anchored in a published academic design, not just a press release.

The gaps are real, and DeepSeek knows it

Semiconductor fabrication facility producing AI accelerators
The real battle over AI hardware is increasingly fought in software: who writes the code decides which chips can actually compete. (Illustration: Calder Brief)

A clear-eyed reading of the release starts with what it does not fix. DeepSeek's own infrastructure remains split: Huawei hardware for inference at scale, Nvidia hardware for training, the most compute-intensive stage of model building. When the lab tried training an earlier model on Huawei silicon, the effort stalled on persistent technical difficulties, according to the Financial Times. Training on domestic silicon is the critical unproven link.

The numbers on that gap are sobering. Liang Wenfeng himself reportedly estimated that four Huawei GPUs equal one Nvidia GPU in effective compute, with Huawei about two years behind. SemiAnalysis found in its August 2026 AgentX benchmark that the CUDA moat remains decisive for multi-step agent tasks, and Huawei's Ascend chips were not even included. The 160,000 Ascend chips DeepSeek has reportedly ordered for its gigawatt-scale data center in Ulanqab, Inner Mongolia, are earmarked for inference, not training. Even on the new supernode, DeepEP-Ascend reached roughly 90 to 95 percent of the physical payload bandwidth limit in tests with up to 32 participating ranks, with larger configurations still being optimized.

None of this is fatal to the thesis. It is the normal shape of an ecosystem challenge: promising at the edges, unproven at the core, improving with every contribution.

Why it matters outside China

Here is the part that travels. DeepSeek did not keep this toolkit as a private competitive advantage. It open-sourced it, and TileLang's other backends mean the flywheel is not China-specific. Any AI developer anywhere who switches from raw CUDA to TileLang on Nvidia hardware is simultaneously lowering the cost of a future hardware switch, to Ascend, to AMD's ROCm, or to Apple's Metal. Each adopter who tests the tool, contributes a kernel, or publishes a benchmark makes it better for the next.

That is precisely how Nvidia won the first round. CUDA was not the best possible language on day one. It was simpler than the alternatives for compute work, and the community that formed around it accumulated the libraries and know-how that made switching away progressively more expensive. The September 30 release suggests the effort has moved from planning to execution.

The inference market is where this lands first. Training demands the bleeding edge; inference, which consumes a growing share of global AI compute, is more forgiving and price-sensitive. If TileLang makes Ascend competitive for serving models at scale, the procurement economics outside the Nvidia ecosystem change everywhere, not just in Hangzhou.

DeepSeek's gamble is that software ecosystems compound. Nvidia spent two decades proving that favors the incumbent; now someone is testing whether it can favor the insurgent too. The code is free, the parity is the pitch, and the clock on CUDA's monopoly just started ticking a little louder.