Hybrid Attention Architecture Emerges as New Standard

Zhipu AI recently disclosed GLM-5.3-Flash project configurations, revealing a 45-layer model architecture combining dual attention mechanisms: 34 layers use Kimi’s KDA linear attention, while 11 layers adopt DeepSeek’s KPool-DSA sparse attention. This configuration signals a shift from early attempts to identify a single efficient attention mechanism toward hybrid designs where multiple mechanisms collaborate with specialized roles.
Concurrently, Qwen3.8-Flash-Next follows a similar approach—maintaining a 3:1 structure of Gated DeltaNet to Attention layers, but replacing the attention layer with Qwen Sparse Attention. MiniMax’s M3 model meanwhile adopted its proprietary MiniMax Sparse Attention (MSA), reducing per-token computation to 1/20 of its predecessor (at 1 million token context). Together, these updates highlight an industry-wide trend: the global attention mechanism, once serving as a safety net, is increasingly being superseded by finely coordinated hybrid alternatives.
From Competition to Collaboration

Attention mechanisms broadly fall into three categories: Full Attention preserves complete token history at high cost; Linear Attention compresses history to reduce computation; Sparse Attention retains token-level history but selectively participates. Historically viewed as competing options, these mechanisms now show complementary potential.
A key counterintuitive development is that Kimi and DeepSeek’s technologies—originally positioned as rival approaches—are now deployed within the same architecture. FLA community maintainer Yang Songlin proposed combining Linear and Sparse approaches as early as 2024, noting KDA and DSA together could cover complementary functions: one maintaining overall state across extended history, the other enabling precise retrieval of specific content.
Technology Diffusion Accelerates Architectural Evolution

Open-source communities, code sharing, and researcher mobility have compressed technology iteration cycles. The FLA open-source community standardized interfaces, kernels, and implementations alongside algorithms, enabling Kimi to build KDA atop shared infrastructure. As techniques escape the confines of academic papers and flow via reusable code across organizations, “technical路线” loses meaning as a static label.
The competitive edge now resides not in possessing specific technologies, but in timing, compositional strategy, and deliberate tradeoffs among capability, cost, and risk. MiniMax accepted Full Attention’s high cost at M2 stage for capability assurance; by M3, Agent workloads made that “insurance” prohibitively expensive, prompting another pivot to sparse attention.
| Model | Attention Architecture | Key Parameters & Context | Primary Driver |
|---|---|---|---|
| GLM-5.3-Flash | 34-layer KDA + 11-layer KPool-DSA | 45-layer hybrid design | Functional specialization: state maintenance vs. precise retrieval |
| Qwen3.8-Flash-Next | 3-layer GDN + 1-layer QSA | Retains previous 3:1 pattern, replaces attention with sparse variant | Global attention functionality migrated to efficient mechanism |
| MiniMax M3 | MiniMax Sparse Attention | Per-token compute reduced to 1/20 of M2 | Agent workflows causing unsustainable scaling of long-context costs |
Implementation Guidance for Practitioners

Teams building new models with hybrid attention are well-advised to adopt KDA+DSA-like patterns when long-context processing and cost tradeoffs coexist with mixed tasks (both reasoning and retrieval). The two-stage division—“state maintenance followed by hotspot retrieval”—mirrors the architectural insights validated in recent releases.
Teams should delay adoption if their applications feature medium or short contexts (<100K tokens); global or lightweight linear attention remains more economical. The engineering complexity of hybrid designs may outweigh theoretical gains in constrained scenarios.
In Conclusion
Attention mechanisms are transitioning from “finding the optimal single solution” toward “defining the optimal composition.” Technology diffusion and compositional innovation have become the engines of large model evolution. What remains uncontested is a team’s ability to preemptively identify the next constraint shaping architectural decisions.