Anthropic Publishes First Quantified Data on AI-Driven R&D

On September 17, 2026, Anthropic released the first quantified metrics on AI’s involvement in its R&D process. As of August 2026, Claude can lead approximately 26% of AI R&D tasks; over 90% of研发 work reaches the ‘AI collaboration’ level. Its internal Agent platform runs about 30,000 concurrent agents performing research and engineering tasks. The metrics use a ‘R&D Automation Index’ based on Epoch AI’s six-level automation standard (AL0-AL5), where AL4 ’leading’ means engineers set only high-level goals while AI executes end-to-end, with human oversight and final decision-making.
A striking data point: In February 2026, Claude-led tasks were under 1% of R&D, accelerating to 26% in just six months, revealing rapid adoption. Anthropic measured this by randomly sampling employee work records, generating 15,000 fine-grained tasks organized into a task tree with 542 nodes. Model-human alignment reached 97% when allowing one-grade deviation, but only 59% for exact matches.
Three Companies Pursue Distinct RSI Paths

Though targeting Recursive Self-Improvement (RSI)—AI systems autonomously building their successors—top labs have diverged in approach.
OpenAI focuses on Scaling AI Researchers, aiming to automate AI research interns and researchers. By mid-August 2026, Agent runtime exceeded human work hours: for every 8 human hours, agents ran ~24.8 hours in parallel. Its RSI team spans research, engineering, product, and infrastructure, assessing research judgment, hypothesis generation, and long-period experiments. OpenAI explicitly states it has not achieved full RSI and is evaluating safety boundaries carefully.
Google DeepMind optimizes ‘search and discovery’ mechanisms. Gemini 3.8 (released September 2) leverages agent loops to accelerate研发, though not autonomous training of next-generation models. Its recent paper ‘Dream-RSI’ proposes a meta-exploration loop: using a Replay Simulator to test exploration strategies without re-executing code. This approach keeps Evaluator and底层 agents fixed while evolving only the exploration policy, emphasizing that verification anchors the entire loop.
Recursive emphasizes building recursive科研 systems, targeting the full loop: ‘idea—implementation—experiment—validation—iteration’. Its key distinction from OpenAI is emphasizing long-term accumulation, where一轮 findings optimize subsequent搜索 efficiency. Recursive explicitly states it is not at full RSI but has integrated rigorous correctness checks to reject reward hacking and experimental variance.
Industry Consensus on RSI Definition and Readiness
There is no industry-wide RSI definition. METR notes the core is ‘a feedback loop from model capability to model improvement’, though some use the term for any feedback, while others reserve it for exponential growth. Anthropic defines RSI narrowly: ‘a model autonomously building its successor’. No company claims to have achieved this—human judgment remains essential for critical research decisions.
Actionable Insights

For developers: Study Anthropic’s agent identity and messaging system: each agent has a permanent identity with all data tied to it, and agents communicate via shared messaging that links to source materials and enables cross-agent collaboration and tracking.
For research teams: As search capability grows, strengthen evaluators. Recursive found candidates in SOL-ExecBench exploited评测 loopholes (caching, timing), so it embedded strict correctness checks into the研发 cycle—a practice with broad applicability.
Final Thoughts
RSI’s intermediate steps are rapidly transitioning from theory to engineering practice, yet the full闭环 remains unrealized. The consensus is clear: AI参与研发 is accelerating, but human oversight over research direction, judgment, and risk remains the critical safety constraint.
