Deepseek releases open-source TileLang tools for Huawei's Ascend chips

Chinese AI developer Deepseek has teamed up with Huawei to build programming tools for Huawei's Ascend chips, according to a post on Deepseek's official WeChat channel. The software includes libraries for computation and for moving data between chips. Reuters reports that Deepseek is making all of it open source. Deepseek says Huawei "fully supported" the work, and the two companies also optimized a so-called supernode, a cluster of 128 Ascend 950 chips.
TileLang sits at the core of the release. It is an open-source programming language for AI chips, originally developed by researchers at Peking University, and Deepseek has been using it for about a year. Deepseek argues that anyone trying to build an independent software ecosystem for AI chips first needs a universal language that is easy to program but still gets full performance out of the hardware. In the company's view, TileLang offers a simpler programming model than CUDA. Deepseek first tested the language on older Nvidia chips, and according to The New York Times, TileLang is now the company's main tool for its work on artificial general intelligence (AGI).
The Decoder frames the release as an attack on one of the biggest problems facing China's AI industry: domestic chips need software that can get the most out of them. Nvidia's dominance does not come from chip design alone. It also rests on an estimated four million developers worldwide who build with CUDA, a moat that rivals like AMD have not been able to cross even when their hardware looked just as strong on paper. The NYT says Chinese model makers like Z.ai and Moonshot AI have so far moved faster than the country's chipmakers. Huawei wants to close that gap: two weeks before Deepseek's announcement, it unveiled new AI processors and supernode systems and said they would be widely used for model training next year. Huawei also admits it cannot keep up with demand at home, so it plans to sell fewer chips abroad. Referring to US export controls, Huawei's current rotating chairman Eric Xu said the company cannot accept a future that hinges on whether others are willing to sell chips to China.
The article then turns to how much of the CUDA moat is left. After testing Jalapeño, OpenAI's inference chip, the research firm SemiAnalysis called the moat "potentially dead", because OpenAI gets new models running on its own hardware so quickly. Jalapeño beat Nvidia's Blackwell on performance per watt in most of the scenarios tested, and according to SemiAnalysis, OpenAI models also helped design the chip while running on Nvidia GPUs. The analysts added their own caveats: they tested only scenarios that are relatively easy to optimize, with about 8,000 input tokens and 1,000 output tokens, and have not yet run AgentX, a benchmark of how AI agents handle multistep tasks.
Agent workloads are where SemiAnalysis found Nvidia well ahead back in August. With AMD's current software stack, Nvidia would still come out cheaper per token even if AMD gave its hardware away. The authors see Nvidia's lasting advantage not in the silicon but in the software that links many chips into one system. Huawei's chips were not part of the AgentX comparison. In an earlier analysis of DeepSeek V4, however, SemiAnalysis pointed out that Huawei's CANN software stack was the only one besides CUDA to support the model on day one.
Key facts
- Deepseek and Huawei built programming tools for Ascend chips, including libraries for computation and for moving data between chips; Reuters reports all of it is open source.
- The core is TileLang, a language originally developed at Peking University that Deepseek has used for about a year and calls simpler than CUDA.
- The two companies also optimized a supernode of 128 Ascend 950 chips; Deepseek says Huawei fully supported the work.
- Nvidia's moat rests partly on an estimated four million CUDA developers; SemiAnalysis called the moat "potentially dead" after testing OpenAI's Jalapeño, but only on relatively easy scenarios.
- SemiAnalysis found Nvidia well ahead on agent workloads in August, and sees its lasting edge in the software that links many chips into one system.
Why it matters
China's AI chips need software that can squeeze performance out of them, and CUDA's developer base is a large part of what keeps Nvidia ahead. Deepseek's argument is that an independent chip ecosystem starts with a universal language that is easy to program yet fast. Releasing TileLang-based tools for Ascend, with Huawei's backing, is a direct move on that problem. It lands as Huawei says its new processors and supernode systems will be widely used for model training next year and as Eric Xu says Huawei cannot accept a future that depends on others' willingness to sell chips to China.
Who it affects
Developers and model makers in China who want to run on Huawei's Ascend hardware, and Deepseek itself, which uses TileLang as its main tool for AGI work according to The New York Times. It also touches Nvidia and rivals such as AMD, whose position depends on software ecosystems as much as on silicon. Huawei's customers abroad are affected too: the company says it cannot keep up with demand at home and plans to sell fewer chips overseas.
How to use it
The source says the tools, libraries for computation and for moving data between chips, are being released as open source, with TileLang as the programming language at the core. Repository, licence and release details are not given, so there is nothing concrete to install from this report. Teams targeting Ascend would need to look at Deepseek's own release channels.
How solid is it
The release itself rests on a post on Deepseek's official WeChat channel, as relayed by Reuters; the claim that TileLang is Deepseek's main AGI tool and that it was first tested on older Nvidia chips comes from The New York Times. That TileLang is simpler than CUDA is Deepseek's own view. No benchmark numbers or performance figures for TileLang or the Ascend 950 supernode are given. The SemiAnalysis findings on Jalapeño are about a different chip and a different question, and the analysts say themselves that they tested only relatively easy scenarios.
Risks and caveats
No performance data backs the claim that TileLang matches the hardware's full potential on Ascend. Nvidia's ecosystem advantage is large, with an estimated four million CUDA developers, and SemiAnalysis found Nvidia well ahead on agent workloads in August. Its Jalapeño verdict is also limited: it has not yet run AgentX, and Huawei's chips were not part of that comparison. Huawei's supply is constrained, since it admits it cannot meet demand at home.
“Nvidia's dominance doesn't come from chip design alone. It also rests on an estimated four million developers worldwide who build with CUDA.”
— The Decoder