A Theory of Continuous Evolution: From End-to-End to Self-Evolution

Mu Yao, a robotics researcher with HKU MMLab background, founded SeeAct AI in 2026 to focus on embodied self-evolution—a direction that traces back to his 2018 graduate work at Tsinghua advocating end-to-end policies and reward-driven learning. Mu currently serves as an tenure-track assistant professor at Shanghai Jiao Tong University, was selected for a national-level young talent program, and has over 5,000 Google Scholar citations. His RoboTwin platform has become a de facto standard in China’s embodied AI community, accumulating over 2,900 GitHub stars, with companies like Ant Group’s LingBot-VA and ShengShu Tech’s Motubrain competing to achieve new benchmarks on it.
Core thesis: the endpoint of embodied intelligence will inevitably be reward-driven self-evolution. According to Mu, robots lack not data but experience. Models trained solely on human demonstration data learn “what humans would do,” but miss the crucial experience of repeatedly encountering failures and adjusting strategies—experience that must be acquired through autonomous interaction with the environment.
Experience Scaling: The True Lever for Scale in Embodied Intelligence
Mu proposes three parallel dimensions for scaling embodied intelligence:
- Data Scaling: expanding training dataset size
- Model Scaling: expanding model parameter count
- Experience Scaling: scaling experience—the dimension most inadequately addressed so far
Experience Scaling encompasses three dimensions:
- Growth in interaction count
- Expansion of experience coverage across scenarios and tasks
- Deepening of experience’s actual value for capability advancement
The core lies in building mechanisms that continuously generate valuable experience: robots encountering new problems can autonomously explore, record trial-and-error outcomes, and reuse them for future learning—while steadily tackling increasingly complex tasks as capabilities improve.
Recursive Self-Improvement: A Snowball Effect of Capability and Experience

Mu introduces “recursive self-improvement” (Recursive Self-Improvement): as capabilities upgrade, systems generate higher-quality experience, which in turn fuels further capability growth—a self-reinforcing cycle.
Take bottle-cap tightening as an example. The full autonomous error-correction process includes:
- Pre-training to accumulate experience with friction variations, grasp misalignment, and other realistic challenges robots may encounter
- Post-training autonomously trying alternatives when facing unseen objects (e.g., opening a handleless black fridge: “try left first, if failed try right”)
- Knowing when to cease ineffective efforts (e.g., not applying 100N force indefinitely to avoid breaking the device)
Critical insight: gains from each interaction must be retained and applied to subsequent tasks—this is what distinguishes true “evolution” from mere data-driven learning.
Architectural Evolution: From RoboTwin to dVLA-RL
Mu’s research follows a clear trajectory:
- RoboTwin 1.0 (originated 2021): lightweight simulation engine enabling full pipelines on student-grade laptops
- RoboTwin 2.0 (2023+): expanded to ~50 tasks, enhanced behavior and code generation, positioned as Benchmark + data generator
- RoboTwin 3.0 (2025+): partnership with Moore Threads for domestic GPU and physical engine support, emphasizing open ecosystem
- VLA & Agent System: shifted from code generation toward end-to-end policy foundation models, culminating in the upper-layer Agent reasoning + lower-layer end-to-end Skill calling architecture (e.g., RoboClaw with AgiHuman Robotics)
- dVLA-RL: multi-task reinforcement learning framework built on discrete diffusion VLA, addressing generalization degradation when switching from single-scenario fine-tuning to new scenarios
On model architecture, Mu’s team found discrete diffusion offers both strong action trajectory generation and policy optimization simplicity—its token prediction mechanism directly suits PPO, GRPO, and other RL methods, making it more practical for reward-driven learning than continuous diffusion/flow-matching approaches.
Deployment Strategy: Virtual-Real Synergy for Experience Scaling

Given high real-world trial-and-error costs, Mu advocates virtual-real synergy:
- True-world, simulation engine, and World Model data constitute three “parallel universes” of experience data
- Dynamic data ratios: determined by current model weaknesses (e.g., background generalization issues call for simulated augmentation; high-precision contact tasks demand true-robot feedback)
- Deployment phase: heavy simulation for exploration, true robots only for critical physics validation and out-of-distribution testing
Notable Contradictory Insight
Though the industry repeatedly cited “data scarcity” in 2024-2025 reports, Mu observes that once appropriate data collection paradigms (e.g., UMI, handheld gripper setups) are found, scaling to millions of hours proceeds faster than expected. The new bottleneck has shifted from data acquisition to experience-generation mechanisms—the central challenge of Experience Scaling and Self-Evolution.
Reader Recommendations

Target audiences:
- Embodied AI algorithm engineers: RoboTwin offers ready-made benchmark infrastructure
- Enterprises planning long-term autonomous robots: Mu’s reward-driven + recursive evolution approach requires early investment in foundation models and simulation ecosystems
- RL practitioners: dVLA-RL’s multi-task joint training framework offers practical reference
Best to wait:
- Industrial applications demanding plug-and-play reliability: current systems still face Sim-to-Real gap and error-correction reliability challenges, far from industrial-grade 99.99% success rates
Final Notes
Mu’s theoretical framework moves embodied intelligence from “data-driven” toward “experience-driven”. Its recursive self-improvement architecture provides an actionable path toward long-term autonomous capabilities. As generative models’ general capabilities become increasingly accessible, enabling robots to truly “grow” self-evolution capacities may become the next competitive frontier.
