Featured image of post Xiaomi Announces MiMo-V3 with New HySparse 2 Architecture, Open-Sources V2.6 Series with Enhanced RL Training Infrastructure

Xiaomi Announces MiMo-V3 with New HySparse 2 Architecture, Open-Sources V2.6 Series with Enhanced RL Training Infrastructure

MiMo-V3 to adopt HySparse 2 architecture; V2.6 series open-sourced with 1M context and 1568-sample batch.

Release Summary: MiMo-V2.6 Open Source and MiMo-V3 Architecture Preview

Release Summary: MiMo-V2.6 Open Source and MiMo-V3 Architecture Preview
Release Summary: MiMo-V2.6 Open Source and MiMo-V3 Architecture Preview|News screenshot

Xiaomi officially released and open-sourced the MiMo-V2.6 series on September 22, while announcing that the upcoming MiMo-V3 will adopt the new HySparse 2 architecture. Key facts:

  • Release date: September 22, 2026 (public announcement on September 23)
  • New versions: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Distill-Qwen-9B
  • Upcoming: MiMo-V3 with HySparse 2 core architecture
  • Pricing: Flash: $1/M input / $2/M output; Pro: $3/M input / $6/M output; 99% caching discount
  • Availability: MiMo-V2.6 live on Xiaomi MiMo open platform; MiMo-V3 no release date yet
  • Weights open: Full weights and technical reports for MiMo-V2.6-Pro and Flash are open-sourced

Breakthroughs in RL Training Scale

Breakthroughs in RL Training Scale
Breakthroughs in RL Training Scale|News screenshot

MiMo-V2.6 represents Xiaomi’s key step toward recursive self-improvement (RSI). Its most significant advancement lies in RL training infrastructure:

  • Single update: 1568 samples, 1M context length training, 3.5-3.7B tokens per step
  • Training time: Under 6 days, cost: ~$2.62M (Pro) and ~$850K (Flash)
  • Total training trajectories: ~750K
  • Task pass rate improvement: +25% (Pro) and +12% (Flash)

RL compute scaling targets three dimensions: larger batch/higher throughput, more complex tasks/agents, and larger Grader capacity. Training environments cover Code, General, Visual, and Cyber categories.

A notable surprise: though Pro scores 46 on AA index v4.3.2 (highest among open-weight models), it still trails closed-source leaders (Claude Fable 5.1, GPT-6 Astra at 53). Even more striking, the base MiMo-V2.6 model (without any Lean-specific fine-tuning)协助 researchers completed a fully verified Lean 4 formal proof of Li-Yorke’s “period three implies chaos” theorem, totaling 6000+ lines demonstrating cross-domain generalization capability.

Capability Comparison and Desktop App

FeatureMiMo-V2.6-ProMiMo-V2.6-Flash
AA v4.3.2 score46Undisclosed
Training cost~$2.62M~$850K
DeepSWE v1.1 delta+17 pts (48.8→65.7)+14 pts (58.4→72.6)
Tokens/step3.5-3.7B3.5-3.7B
SWE-bench Verified61.1→66.2Undisclosed
API pricing ($/M)3/61/2

The MiMo Desktop client (Windows/macOS) launched with UltraSpeed mode offering up to 20x inference speedup. Subscribers gain direct access to Pro/Flash; API Key users configure custom endpoints.

Open-Source Contributions and Research

Open-Source Contributions and Research
Open-Source Contributions and Research|News screenshot

Three core assets open-sourced:

  1. 7k+ RL task environments: Software engineering, vulnerability reproduction, knowledge work, web dev
  2. End-to-end RL training framework: Built on verl, uni-agent, mini-swe-agent
  3. Lightweight modular harness: mini-harnesses for Multi-Harness Training improves domain generalization

Research use cases: MOF material design screening and formal proof (Lean 4) of Li-Yorke “period three implies chaos” theorem, fully verified by Lean kernel.

Adoption Recommendations

Adoption Recommendations
Adoption Recommendations|News screenshot

  • Try now if: You need 1M context processing; cost-sensitive API deployment; interested in RL framework reproduction or multi-agent systems
  • Wait if: Your production system requires <50ms latency with strict SLA; you need AA >50 scores and cannot accept open-weight weights (wait for V3 launch)

Final Notes

Xiaomi’s open-sourcing of 1M-context, 1568-sample-batch RL training infrastructure provides Chinese developers rare access to industrial-grade reinforcement learning tooling. Combined with price points 1/20-1/60 of overseas alternatives, this may accelerate open-weight models’ adoption in long-context and agentic computing.