Release Summary: MiMo-V2.6 Open Source and MiMo-V3 Architecture Preview

Xiaomi officially released and open-sourced the MiMo-V2.6 series on September 22, while announcing that the upcoming MiMo-V3 will adopt the new HySparse 2 architecture. Key facts:
- Release date: September 22, 2026 (public announcement on September 23)
- New versions: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Distill-Qwen-9B
- Upcoming: MiMo-V3 with HySparse 2 core architecture
- Pricing: Flash: $1/M input / $2/M output; Pro: $3/M input / $6/M output; 99% caching discount
- Availability: MiMo-V2.6 live on Xiaomi MiMo open platform; MiMo-V3 no release date yet
- Weights open: Full weights and technical reports for MiMo-V2.6-Pro and Flash are open-sourced
Breakthroughs in RL Training Scale

MiMo-V2.6 represents Xiaomi’s key step toward recursive self-improvement (RSI). Its most significant advancement lies in RL training infrastructure:
- Single update: 1568 samples, 1M context length training, 3.5-3.7B tokens per step
- Training time: Under 6 days, cost: ~$2.62M (Pro) and ~$850K (Flash)
- Total training trajectories: ~750K
- Task pass rate improvement: +25% (Pro) and +12% (Flash)
RL compute scaling targets three dimensions: larger batch/higher throughput, more complex tasks/agents, and larger Grader capacity. Training environments cover Code, General, Visual, and Cyber categories.
A notable surprise: though Pro scores 46 on AA index v4.3.2 (highest among open-weight models), it still trails closed-source leaders (Claude Fable 5.1, GPT-6 Astra at 53). Even more striking, the base MiMo-V2.6 model (without any Lean-specific fine-tuning)协助 researchers completed a fully verified Lean 4 formal proof of Li-Yorke’s “period three implies chaos” theorem, totaling 6000+ lines demonstrating cross-domain generalization capability.
Capability Comparison and Desktop App
| Feature | MiMo-V2.6-Pro | MiMo-V2.6-Flash |
|---|---|---|
| AA v4.3.2 score | 46 | Undisclosed |
| Training cost | ~$2.62M | ~$850K |
| DeepSWE v1.1 delta | +17 pts (48.8→65.7) | +14 pts (58.4→72.6) |
| Tokens/step | 3.5-3.7B | 3.5-3.7B |
| SWE-bench Verified | 61.1→66.2 | Undisclosed |
| API pricing ($/M) | 3/6 | 1/2 |
The MiMo Desktop client (Windows/macOS) launched with UltraSpeed mode offering up to 20x inference speedup. Subscribers gain direct access to Pro/Flash; API Key users configure custom endpoints.
Open-Source Contributions and Research

Three core assets open-sourced:
- 7k+ RL task environments: Software engineering, vulnerability reproduction, knowledge work, web dev
- End-to-end RL training framework: Built on verl, uni-agent, mini-swe-agent
- Lightweight modular harness: mini-harnesses for Multi-Harness Training improves domain generalization
Research use cases: MOF material design screening and formal proof (Lean 4) of Li-Yorke “period three implies chaos” theorem, fully verified by Lean kernel.
Adoption Recommendations

- Try now if: You need 1M context processing; cost-sensitive API deployment; interested in RL framework reproduction or multi-agent systems
- Wait if: Your production system requires <50ms latency with strict SLA; you need AA >50 scores and cannot accept open-weight weights (wait for V3 launch)
Final Notes
Xiaomi’s open-sourcing of 1M-context, 1568-sample-batch RL training infrastructure provides Chinese developers rare access to industrial-grade reinforcement learning tooling. Combined with price points 1/20-1/60 of overseas alternatives, this may accelerate open-weight models’ adoption in long-context and agentic computing.
