Tongyi Lab Unveils Full-Stack AI Infrastructure with Wuzhen V900 Chip
- Launch Date: September 22, 2026 at Hangzhou Yunting AI Conference main stage
- New Versions: Wuzhen V900 AI chip, ICN Switch interconnect chip, Panmai smart NIC, Zhenyue SSD controller, Yitian CPU roadmap
- Availability: V900 expected in mass production Q1 2027; Wuzhen J900 planned for Q3 2028
- WeightOpen Status: Not disclosed; SAIL software stack supports 260+ frameworks
- Target Workloads: Large model training/inference, embodied intelligence, autonomous driving, finance, energy, manufacturing
Wuzhen V900: 3x Performance, Doubled VRAM, Enabling Trillion-Parameter Models

Tongyi Lab’s latest Wuzhen V900 AI chip achieves triple performance over its predecessor M890, with 216GB VRAM (50% increase from 144GB), and 1200GB/s inter-chip bandwidth (up from 800GB/s). It natively supports FP8 and FP4 precision for both high-accuracy training and low-precision inference.
With 216GB VRAM, deploying trillion-parameter models requires only 10+ V900 chips. Baseboards using M890 have already run Qwen3.8 and Kimi K3 models exceeding 2 trillion parameters. A key surprise: As models scale to multi-trillion parameters, even distributed training runs into VRAM limits per chip—此时 the interconnect speed between chips, not raw compute, becomes the bottleneck for expert parallelism efficiency.
Full-Stack Synergy: From Chip to Server System Optimization

V900’s full potential emerges through Tongyi Lab’s integrated stack:
- Wuzhen V900: Core AI compute accelerator
- ICN Switch: Enables thousand-node high-bandwidth interconnect with native memory semantics and unified addressing
- Panmai Smart NIC: Connects servers to datacenter networks
- Zhenyue SSD Controller: Enhances local storage throughput
- Yitian CPU Roadmap: 2027’s Yitian 720 (190 cores, 12-channel GDI) for throughput-focused workloads; Yitian 730 as first self-developed高性能 core with 1.4x single-core performance over predecessor for latency-sensitive AI Head Node tasks
For Agentic workloads, CPU-AI chip collaboration intensifies. Intel presented data showing at this year’s Computex that CPU:GPU ratios are shifting from 1:8 in training to 1:1 or higher in Agent scenarios.
Hardware + Software Co-Design: Accelerating Model Deployment

SAIL software stack provides OS coverage, SDKs, and debugging tools. Real-world impact: Kimi K3 adaptation on M890 achieved 35% first-token latency reduction and 1.8x decoding throughput through software optimization. SAIL now supports 260+ frameworks, covering 96% of top AI repositories (15,218 repos with 100M+ stars), with vLLM/SGLang adapters typically deployed within one week.
Alibaba Cloud’s Lingjun Wuzhen super-node instances deliver up to 1.5x performance gain in Agentic inference scenarios, validating end-to-end chain optimization from chip to cloud.
| Chip Version | VRAM | Inter-Chip Bandwidth | Performance vs. Predecessor | Release Date |
|---|---|---|---|---|
| Wuzhen 810E | Not disclosed | — | — | 2024 |
| Wuzhen M890 | 144GB | 800GB/s | 3x on 810E | May 2026 |
| Wuzhen V900 | 216GB | 1200GB/s | 3x on M890 | Sept 2026 (mass prod. Q1 2027) |
| Wuzhen J900 | Not disclosed | Not disclosed | Not disclosed | Q3 2028 |
Target Audience & Deployment Advice

- Early adopters: Enterprises deploying trillion-parameter MoE models, especially in embodied intelligence and autonomous driving where training efficiency and cluster stability are critical. Existing M890 users (e.g., Zhongqing, XPeng) can evaluate V900’s performance leap.
- Wait: Companies with modest inference needs or non-Agent workloads should wait for V900 pricing and ecosystem maturity; M890 remains cost-effective for steady-state inference.
Final Thoughts
AI hardware is shifting from raw chip benchmarks to system-level efficiency wars. Alibaba’s integrated stack—including models, chips, and cloud—positions it uniquely to navigate post-Moore’s Law compute limitations through hardware-software co-design.
