Stepfun Launches Step 5 Preview: MoE Flagship with Global Top-3 Ranking

Stepfun officially released its new MoE (Mixture of Experts) flagship model, Step 5 Preview, on September 20, 2026. Designed for real-world Agentic tasks, the model features an impressive 600 billion total parameters with only 27 billion activated per inference. Key facts:
- Release date: September 20, 2026
- Version: Step 5 Preview (preview edition)
- Pricing: Input (cache miss) ¥7/M token, Input (cache hit) ¥0.35/M token, Output ¥20/M token
- Availability: API access live now; open source on October 15, 2026
- Weight release: Preview available via API; full开源 release planned for mid-October
Core specs: 600B total parameters, 27B activated, 1M token context window, native multimodal (text + vision) support.
Surprising Performance: Open-Source #3, Cost at 1/8 of Top Commercial Models

Step 5 Preview scored 44 points in the Artificial Analysis Intelligence Index, ranking third among open-source models globally—trailing only Qwen3.8 Max (0902) and GLM-5.3 (max). Its output speed of 100 tokens/second places sixth on the same leaderboard.
A striking contrast emerges in cost-performance: the model performs just below GPT-6 Astra and Claude Opus 5 on benchmark tests, yet operates at one-eighth the cost—$0.71 (¥4.76) per task versus Claude Opus 5 at $5.86 (¥39.25).
In internal benchmarks covering long-horizon CLI tasks, financial research (FrontierFinance), and cross-domain analysis (DRACO), Step 5 Preview outperforms Kimi K3 and GLM-5.3, and scores 49 on StepCodeBench (553 code repos, 20 domains). A particularly notable result: optimizing a GPU kernel took 22 hours on H100, achieving 508 TFLOPS—surpassing Claude Opus 5’s 493 TFLOPS**.
Multi-Scenario Testing: From PPT to 3D Apps and Game-Playing Agents

Real-world testing reveals strong capabilities in complex task orchestration:
- Content conversion: Turns blog posts into well-structured PPTs (though visual variety remains limited)
- Dynamic web generation: Maps abstract descriptions (e.g., “a flowing poem”) to interactive visual pages
- 3D development: Uses Three.js to build interactive hot air balloon scenes with real-time parameter control
- Business simulation: Completes 7-stage tea shop operation simulations with cash flow projections and strategy summaries
Engineering tasks include automatic ESP32-S3 firmware改造 and 6M-token continuous gameplay in Pokémon Red (3000+ turns) without task-specific tuning—demonstrating robust long-horizon planning.
| Model | Total Params | Active Params | Cost per Task (USD) | Cost per Task (CNY) | Art. Index Score | Open Source |
|---|---|---|---|---|---|---|
| Step 5 Preview | 600B | 27B | 0.71 | 4.76 | 44 | Oct 15, 2026 |
| Claude Opus 5 (max) | Not disclosed | Not disclosed | 5.86 | 39.25 | >44 | Closed |
| Kimi K3 | Not disclosed | Not disclosed | Not disclosed | Not disclosed | <44 | Not disclosed |
| GLM-5.3 (max) | Not disclosed | Not disclosed | Not disclosed | Not disclosed | >44 | Partial |
| Qwen3.8 Max (0902) | Not disclosed | Not disclosed | Not disclosed | Not disclosed | >44 | Partial |
Who Should Adopt Now—and Who Should Wait

Ready to try now if you:
- Process long-context tasks (≥100K tokens) like document analysis, workflow automation, or multi-turn CLI orchestration
- Need cost-efficient Agent deployment for research, education simulation, or operations
- Seek open-source alternatives to reduce long-term API spend
Wait if your use case:
- Demands rich visual diversity without frequent human iteration (design-heavy workflows)
- Requires enterprise SLAs or strict compliance certifications not yet available
- Operates in closed hardware ecosystems lacking API exposure
Final Thoughts
Step 5 Preview exemplifies the industry shift from raw throughput to computational efficiency: the MoE architecture delivers 600B-parameter capability at 27B activation cost, making intelligence more deployable. As AI models move from benchmarks to production, efficiency—not just scale—will define the winners of the Agent era.
