DeepSeek Releases Open-Source Software Stack for Huawei AI Chips
DeepSeek has launched an open-source programming toolkit designed to support Huawei’s Ascend AI accelerators, providing a domestic alternative to Nvidia’s CUDA ecosystem.

DeepSeek announced on September 30 that it has open-sourced a comprehensive suite of programming tools built for Huawei Technologies’ Ascend AI accelerators. The release aims to provide Chinese developers with a homegrown alternative to Nvidia’s CUDA framework, which has long dominated AI chip programming.
The toolkit features TileLang, a high-level programming language intended to function as a domestic counterpart to CUDA. By offering a similar abstraction layer, the software allows developers familiar with Nvidia’s workflow to transition to Huawei hardware without needing to learn an entirely new system.
Beyond TileLang, the package includes several compute and communication libraries such as DeepGEMM, FlashMLA, TileKernel, DeepSelect, and DeepEP. These tools are designed to manage various stages of the AI development process, including matrix multiplication, memory management, and inter-chip communication.
The software is optimized for Huawei’s Ascend 950 chips and supports a supernode configuration that links 128 accelerators together. DeepSeek made the tools available for free download to lower the barrier for developers who might otherwise rely on Nvidia due to high switching costs.
This release addresses a strategic challenge for China following US export controls on advanced Nvidia chips. While companies like Huawei have designed domestic silicon, the lack of a mature software ecosystem previously limited the utility of that hardware. DeepSeek received support from Huawei throughout the development process.
To further support this infrastructure, DeepSeek plans to build a data center in Inner Mongolia capable of hosting at least 160,000 Huawei Ascend accelerators. This follows the company's efforts in April 2026 to adapt its V4 AI model to run on Huawei hardware, demonstrating that high-end models can function on non-Nvidia silicon.
The toolkit also includes benchmarking utilities to help developers measure performance and optimize code for the Ascend architecture.



