On September 30, 2026, DeepSeek and Huawei were reported to be collaborating on software for Huawei’s Ascend AI processors. The same coverage described a 128-chip Ascend 950 supernode. The tools are meant to help developers write AI kernels—small programs that carry out computations on a processor—but the reported chip count is not a performance result.

What software was reported for Ascend?

The reported software lineup included Ascend adaptations of TileLang, DeepGEMM, DeepEP and FlashMLA. The work was also described in broader terms as TileLang alongside compute and data-transfer libraries.

TileLang-Ascend, a project for programming Huawei Ascend NPUs, is designed to help developers create AI compute kernels. Its documentation describes two compiler routes: Ascend C with PTO, and AscendNPU IR. Existing CUDA code was reported to require rewriting for the Ascend backend; the software was not described as automatically converting CUDA kernels.

What does TileLang-Ascend document?

The project repository lists Ascend A2 and A3 as tested and validated devices. It records that TileLang-Ascend became open source on September 29, 2025, and lists DeepSeek V4 kernels on April 24, 2026—both before the September 2026 collaboration coverage.

For the installation setup described by the project, the stated prerequisites are CANN 8.3.RC1 or later and torch-npu 2.6.0.RC1 or later. Those requirements and the listed A2 and A3 devices describe the project’s documented setup and test scope.

What does the reported 128-chip Ascend 950 system tell us?

The 128-chip figure describes the scale of the reported Ascend 950 supernode and the system DeepSeek and Huawei were said to have optimized. A chip count measures configuration, not speed. It does not establish how that system performs against an equivalent Nvidia system; that comparison depends on results from comparable workloads and hardware.