Core Announcement: Major Breakthrough in Domestic AI Chip Toolchain

In September 2026, DeepSeek open-sourced its core operator toolkit TileLang with native support for Huawei Ascend 950 NPU. Key details include:
- TileLang, the flagship operator development product, officially supports Ascend platform natively
- Five core libraries—DeepGEMM, DeepEP, FlashMLA, TileKernels, and DeepSelect—are released with Ascend versions
- Unlike traditional CUDA-bound stacks, TileLang uses decoupled architecture for cross-hardware execution
- As of September 2026, TileLang’s GitHub repository accumulated 7.6k stars and 1,992 commits
This marks the first time domestic AI chips achieve functional parity with NVIDIA’s ecosystem at the core operator level.
Technical Architecture: Decoupling Solves Cross-Platform Challenges

TileLang’s key innovation is a decoupled frontend-backend architecture with separated target and execution modules. Traditional compilers tightly bind hardware-specific logic with general compilation workflows, requiring full rebuilds for new chips. TileLang separates into two modules:
- Target backend: Only defines instruction syntax and optimization rules for specific hardware (e.g., WGMMA/TCGEN05 for CUDA, MFMA/WMMA for ROCm, Ascend C vector instructions)
- Execution backend: Handles universal compilation, loading, and runtime workflows, independent of specific hardware
Adapting new chips requires only adding a new target backend; the execution backend remains unchanged. Each backend follows a four-layer pipeline:
- Dialect layer: Translates to chip-specific instruction dialects
- Context layer: Normalizes hardware specifications and tracks compilation state
- Pass pipeline layer: Core optimization generating efficient low-level code
- Codegen layer: Splits host (CPU) and device (NPU/GPU) code
This enables developers to explicitly control memory hierarchy (register/shared/global), thread scheduling (tiling grids), and parallelism (software pipelining) in approximately 30 lines of Python code—compared to 500+ lines in CUDA.
Key Metrics: Performance Parity with Industry Leaders, Broad Hardware Coverage

As of September 2026, TileLang supports mainstream AI platforms: CUDA is the feature-complete flagship; AMD ROCm, Apple Metal, and Ascend 950 are officially supported; LLVM CPU and WebGPU are experimental; MindCraft and Moore Threads use ecosystem adaptations.
The key metric is that on NVIDIA H100, TileLang not only outperforms PyTorch native (2-20x speedup) but also beats OpenAI Triton and matches NVIDIA’s optimized cuBLAS and FlashAttention—proving open tools now reach industry-top performance.
| Feature | TileLang | PyTorch Native | OpenAI Triton | NVIDIA cuBLAS |
|---|---|---|---|---|
| Peak performance | Near hardware limit | Medium | High | Highest (hardware-tuned) |
| Cross-platform | Native multi-chip support | CUDA-only | Partial support | CUDA-only |
| Development complexity | ~30 Python lines | High | Medium | N/A (proprietary) |
| Flexibility | High (custom/fusion) | Low | Medium | None |
Adoption Recommendations: Who Should Act Now?

- Ready to adopt: Teams requiring multi-hardware deployment capabilities; engineering groups seeking lower operator development门槛 and higher productivity; cloud providers and AI labs with Ascend purchases or plans
- Consider waiting: Production users needing extended Ascend validation; developers relying solely on PyTorch APIs without custom ops; enterprises requiring longer track record verification
Final Thoughts
TileLang’s value lies not in replacing CUDA, but in closing the operator development gap for domestic chips. When component-level tools become open and hardware-agnostic, autonomy shifts from hardware alone to a coordinated software-hardware ecosystem.