The Efficiency Battle for Trillion-Parameter Models Intensifies: Inspur Unveils Super-Node Server, Overseas Team Open-Sources K3-Optimized Inference Megakernels

Inspur and overseas teams tackle 2.8T-parameter Kimi K3 inference limits via super-node servers and megakernel optimization.

Core Event: Super-Node and Megakernel Approaches Break Through Trillion-Parameter Inference

Core Event: Super-Node and Megakernel Approaches Break Through Trillion-Parameter Inference
Core Event: Super-Node and Megakernel Approaches Break Through Trillion-Parameter Inference|News screenshot

In late September 2026, domestic and overseas teams nearly simultaneously unveiled inference optimization solutions targeting the world’s largest open-weight model, Kimi K3, marking the efficiency competitionfor large models has entered a deep phase at the system level:

  • Sept 21, 2026: Inspur released the Metacore SD200 Ultra super-node AI server, single-machine deploying the 2.8T-parameter Kimi K3, with token generation latency突破 5.85ms
  • Sept 23, 2026: Overseas inference optimization team Inferact open-sourced tpu-megakernels, using 16 Google TPU v7 chips to fit the entire K3 model within a single kernel
  • Kimi K3 officially recommends 64+ chips for deployment; unoptimized initial performance is ~10 tokens/s
  • Inspur’s measured token latency: 5.85ms, single-user throughput: 170 tokens/s; Inferact’s low-concurrency decode throughput: 709 tokens/s

Operator Fusion: Breaking Kernel Boundaries to Address Memory Bottlenecks

Operator Fusion: Breaking Kernel Boundaries to Address Memory Bottlenecks
Operator Fusion: Breaking Kernel Boundaries to Address Memory Bottlenecks|News screenshot

Kimi K3’s 2.8T total parameters and 104B activated parameters require ~1.5TB HBM bandwidth per forward pass, triggering over 120 All-to-All communications. Traditional inference frameworks suffer"bubble" delays from excessive fragmented kernel launches.

A counterintuitive contrast:

  • Inferact’s overseas team chose extreme fusion: one megakernel for the entireDecode process, leveraging TPU v7’s 64 MiB on-chip VMEM for explicit data lifetime control, overlapping computation with data prefetching
  • Inspur adopted an engineering-pragmatic approach: fusing numerous base operators into super-operators, reducing operator count by 10x, with communication-computation synchronized pipelines delivering 3x+ performance gains

Both paths converge on one truth: the traditional “one Op, one Kernel” loose coupling model has hit its ceiling; system-level fusion has become the new bottleneck breakthrough.

Super-Node System Battle: Three Breakthroughs from 128-Chip Tight Coupling

Inspur’s SD200 Ultra builds system-level advantages with three core innovations:

  1. 3D Hyper Mesh Interconnect: Tight coupling of 128 domestic AI chips, communication latency as low as 0.69μs
  2. Unified Address Space for GPU Memory: 8TB unified GPU memory and 64TB system RAM, with symmetry memory technology enabling the first GPU memory direct-connect on domestic chip super-nodes, achieving 3.5x faster AllReduce
  3. Super-Operator Agent: Kimi K3 itself participates in optimization—the agent generates tile-level pipeline super-operators for the K3 model, delivering 50% further performance boost over manual optimization via a “K3 optimizes K3” feedback loop

Paradigm Shift in Inference Optimization: AI as an Optimization Tool

Paradigm Shift in Inference Optimization: AI as an Optimization Tool
Paradigm Shift in Inference Optimization: AI as an Optimization Tool|News screenshot

Inspur’s “super-operator agent” transforms the model from an optimization target into an optimization tool itself, summarized by the high-speed rail vs. green-slot train analogy:

  • Traditional approach: Hundreds of operators chaining sequentially, repeatedly writing intermediate results to HBM
  • Super-operators: Results directly stay in registers/on-chip caches for next operators, implementing prefetch-compute-writeback as a three-stage concurrent pipeline

This paradigm significantly reduces manual optimization burden, enabling sustainable system-level efficiency iterations.

Adoption Guidance

Adoption Guidance
Adoption Guidance|News screenshot

  • Megakernel users (DSPark+tpu-megakernels): Teams with TPU v7 hardware seeking ultra-low decode latency—open-source tpu-megakernels Enable rapid prototyping
  • SD200 Ultra super-node adopters: Commercial deployments requiring stable 2.8T-parameter model inference with国产化 requirements, prioritizing unified memory addressing and guaranteed system throughput
  • Waiting recommended for: Buyers seeking maximum single-chip performance should monitor domestic chip ground-up capabilities; current super-node solutions still face single-chip performance ceilings

Final Thoughts

As AI compute competition shifts from chip specifications to system organization, Chinese teams have forged a deterministic追赶 path under chip-technology constraints via super-node and super-operator innovations. The convergence of megakernel and super-operator approaches confirms the inference efficiency revolutionhas decisively moved from model layer to system layer deep water.