Featured image of post Inferact Achieves 57% Speedup on TPU over GPU for Kimi K3 Using Megakernel

Inferact Achieves 57% Speedup on TPU over GPU for Kimi K3 Using Megakernel

vLLM founding team releases TPU-specific megakernel, outperforming GPU on Kimi K3 inference by 57%.

Core Event

Core Event
Core Event|News screenshot

Inferact unveiled its TPU optimization results in September 2026: running Kimi K3 on 16 TPU v7 Ironwood chips achieved 709 tokens/s—a 57% speedup over 16 GB200 chips at 452 tokens/s. Results were measured under identical conditions using vLLM and DSpark.

  • Test Platform: 16× TPU v7 Ironwood vs 16× GB200
  • Inference Stack: vLLM + DeepSeek DSpark
  • Models: Kimi K3, Qwen 3.8 27B
  • Status: Code open-sourced (tpu-megakernels repo)
  • Funding: $150M Series A led by a16z, $800M valuation

Megakernel Breakthrough: Eliminating Kernel Boundaries

Megakernel Breakthrough: Eliminating Kernel Boundaries
Megakernel Breakthrough: Eliminating Kernel Boundaries|News screenshot

Founded by vLLM’s original team, Inferact’s innovation is megakernel technology—fusing hundreds of independent small kernels into a single unified program.

Traditional inference suffers from memory bandwidth idling between kernel launches. NVIDIA’s CUDA Graphs and PDL reduce but cannot eliminate this gap. Inferact’s solution uses Google Pallas to hand-code the full 92-layer MoE forward pass as one kernel call.

TPU’s hardware naturally supports this: each TensorCore has 64 MiB VMEM managed entirely in software. Engineers can precisely schedule data loading so that layer N+1 weights begin prefetch while layer N computes—parallelizing computation and data movement.

Megakernel: A technique merging micro-kernels into one unified kernel to eliminate scheduling overhead

Counterintuitive Data: Higher Bandwidth, Slower Performance

A key surprise: TPU v7 offers 7380 GB/s HBM bandwidth versus GB200’s 8000 GB/s—yet the lower-bandwidth TPU outperforms. This confirms software optimization as the performance driver.

Performance Benchmark

Performance Benchmark
Performance Benchmark|News screenshot

SetupModelTPU v7GB200Advantage
16 chipsKimi K3709 token/s452 token/s+57%
16 chipsKimi K3 (no speculative)249 token/s (batch=1)127 token/s+96%
16 chipsKimi K3 (no speculative)865 token/s (batch=8)636 token/s+36%
4 chipsQwen 3.8 27B1515 token/s695 token/s+118%

DSpark speculative decoding achieved 6-token acceptance length on TPU, with ~8.5ms per decode step.

Accuracy Verification & Usage Guidance

Accuracy Verification & Usage Guidance
Accuracy Verification & Usage Guidance|News screenshot

greedy decoding confirmed zero accuracy loss: Kimi K3 scored 94.4% on GPQA-Diamond and 97.2% on GSM8K—identical to GPU results.

  • Best for: Organizations with existing TPU infrastructure or planning deep customization; workloads favoring high batch throughput
  • Wait for now: Smaller teams needing multi-model compatibility (megakernel currently Kimi K3-optimized; re-adaptation required for new architectures)

Final Thoughts

This performance reversal demonstrates a paradigm shift in AI infrastructure: software stack efficiency now outweighs raw hardware specs. As the vLLM founding team drives TPU ecosystem maturity, hybrid approaches combining automated tools like OpenXLA with hand-crafted megakernels may redefine inference framework competition.