Core Announcement
At the AI Infrastructure Summit held on September 16, 2026 in Santa Clara, California, Samsung Electronics disclosed its next-generation AI memory technology, zHBM, with the goal of boosting AI accelerator response speed by 10x. Key milestones:
- Release timeline: zHBM launched in August 2026; zNAND-O 3D storage samples expected from 2028
- Performance target: Increase token processing from 100 tokens/user/second to 1,000 tokens/user/second
- Peak specifications: Up to 8x performance and 3x improvement in energy efficiency over HBM5
- Architectural shift: Vertical stacking of HBM directly atop AI accelerators
Technical Innovation and Industry Skepticism
Samsung’s金仁东 (DSA storage product planning executive) compared zHBM’s vertical stacking design to “installing a dedicated elevator from a hotel room directly to the lobby”—eliminating data transport latency inherent in conventional horizontal layouts.
An unexpected contrast emerged regarding CXL (Compute Express Link), a complementary memory technology to HBM. OpenAI researcher Daniel Morris expressed strong reservations: “For actual AI model execution, I can’t find where CXL can be used.” Its sole validated use case, per Morris, is storing rarely-accessed cold data.
Vidya Thiagarajan, Intel’s AI SoC architecture director, reinforced this view: “While CXL-based memory integration helps, it supplements auxiliary storage and cannot replace HBM. Data transfer between GPUs via CXL is far slower than with HBM.”
Technical Comparison: HBM vs. CXL
| Aspect | HBM (including zHBM) | CXL |
|---|---|---|
| Physical structure | Vertical stacking via TSV (through-silicon vias) adjacent to processor | Horizontal connection via third-party interface |
| Bandwidth capability | zHBM up to 8x HBM5 theoretical peak | Significantly lower, suitable for capacity scaling over bandwidth |
| Primary use case | AI model hot data, real-time inference, weight matrix access | Server memory expansion, cold data storage, cross-server memory sharing |
| Main vendors | Samsung, SK Hynix | Intel-led, broad industry support |
The comparison suggests CXL serves as an auxiliary memory extender while high-concurrency inference demands remain firmly in HBM’s domain.
Practical Recommendations
- Act now if: You design AI chips or operate large-scale inference platforms. For real-time conversational AI (customer service bots, voice assistants), zHBM’s low-latency architecture delivers measurable UX improvements. Monitor Samsung’s 2027-2028量产 timeline.
- Wait longer if: Budget constraints dominate your AI deployment. Current HBM remains costly and supply-constrained. CXL may offer more economical alternatives for non-real-time workloads. Consider zNAND-O alternatives post-2028.
Final Thought
The summit revealed a pivotal industry inflection point: the bottleneck has shifted from compute units to memory subsystems. As Moore’s Law slows, restructuring the memory-processor topology through 3D stacking emerges as the critical path to sustained performance growth—transforming memory vendors like Samsung from component suppliers into architectural gatekeepers.
