In June this year, the Anthropic Institute, an arm of Anthropic, published a long-form article titled "When AI Builds Itself," keeping the entire AI industry on high alert.
The core judgment of the article is straightforward: given sufficient computing power, AI systems could potentially design and develop their own successor models entirely autonomously, a process known as "recursive self-improvement"; this prospect "may arrive sooner than most institutions are prepared for," significantly elevating the risk of humanity losing control over technology.
This is not a private prophecy from some doomsday theorist, but a public warning issued by a top-tier AI laboratory regarding its own field.
01. What RSI Really Is, and What It Is Not
The core logic of Recursive Self-Improvement, abbreviated as RSI, is not overly complex: AI is no longer just a tool that passively executes human instructions, but rather the driver of its own evolution. It can autonomously identify model flaws, write optimized code, and design new architectures. Each iteratively upgraded version is more capable than the previous one, enabling it to complete the next round of self-optimization with higher efficiency, thereby forming a snowballing positive feedback loop of capabilities. However, there is a much wider gap than imagined between "AI modifying code once" and true RSI.
Yuxuan Zhang, a researcher at the University of British Columbia, provided a rather restrained definition in a widely circulated blog post: RSI is not a milestone achieved simply because "AI has modified itself," but rather a strongly conditioned claim—the system has improved the mechanism for executing the next round of improvement, and subsequent generations remain better at continuing to improve even on evaluations they were not selected on.
In other words, true RSI requires demonstrating an improvement capability that is transmissible across generations and still holds true in independent evaluations, rather than just gaming the scores on the same evaluator. Shunyu Yao further clarified this boundary during a discussion at AGI House. In his view, having a model train itself using its own generated data is merely the most basic form of reinforcement learning and does not yet qualify as RSI.
"If the model can design the entire training process on its own, this already goes beyond the scope of RL." In traditional RL, the training scheme is predefined by humans, and the model is only responsible for generating data; whereas RSI requires the system to not only produce results but also rewrite the method for generating results next time. This means RSI is a systems engineering problem from the very beginning: the scheduling layer, context management layer, and hardware optimization layer all need to be fully integrated, and unilateral capability improvements of the model are far from sufficient.
02. Verification Is Harder Than Implementation
An easily overlooked fact is that the most tricky part of RSI lies not in getting AI to modify itself, but in proving that "it actually became stronger after the modification."
In the blog post, Zhang decomposed self-improvement into hierarchical levels: output-level iteration is merely the same model self-revising its answers; prompt and workflow optimization belongs to constrained scaffolding adjustments; policy optimization in a fixed environment is effective RL; one level above that involves humans setting the verifier and resource budget while the system iteratively optimizes a specific program or training script; only when reaching the level of "improving the next generation's model, training process, toolchain, or evaluation process" does it touch the threshold of strong RSI.
Most current experimental results only support the lower levels—iteratively optimizing a specific prompt, program, or workflow under fixed objective functions and verifiers, rather than strong RSI. This introduces a more insidious risk: evaluation overfitting. Citing research by Dwork et al., Zhang pointed out that repeatedly selecting and tuning parameters on the same set of evaluation benchmarks produces "adaptive overfitting," which is a rigorous statistical issue.
During that discussion at AGI House, a guest also bluntly stated that public benchmarks are often gamed, and surviving for three to five months is already quite an achievement. The three verification thresholds proposed by Zhang—whether the object of modification touches the system that generates the next generation, whether the closed evaluation for selecting offspring's development and reporting capabilities is separated, and whether reliable improvements of v2 > v1 > v0 still emerge under equal resources, unseen tasks, and an altered testing framework—serve precisely as a diagnostic framework for this issue.
03. Gauging Real Progress Through Code Proportions
Setting aside conceptual debates, industry data provides a more intuitive observational thread. Internal statistics disclosed by Anthropic in September 2026 show that as of August 2026, Claude had taken the lead in 26% of the company's AI R&D tasks, a figure that was less than 1% in February; as of May 2026, over 80% of the code merged into the repository was generated by Claude.
The proportion disclosed by Microsoft (MSFT.O) in 2025 was 20% to 30%, while a survey released by New Relic in July 2026 showed that 67% of technology leaders estimate that 51% to 75% of the code produced by their organizations weekly comes from AI. These figures all point to the same conclusion: AI's role in the R&D process is shifting from an "auxiliary tool" to a "primary executor."
Anthropic maps this progression onto a timeline: from 2021 to 2023, humans wrote the code; from 2023 to 2025, chatbots assisted; from 2025 to 2026, coding agents began writing and modifying code on their own; today, autonomous agents can run code independently and delegate hours of work to other agents. The next step is the "closed loop"—agents building and training models themselves. As of now, core decision-making power remains in human hands, but the speed at which humans are transitioning from "sole designers" to "supervisors" is faster than most people anticipated.
04. The Divergence Is Not in Technology, but in Pace
The real debate sparked by RSI is not whether it can be achieved, but at what pace it should be realized. On September 21, 2026, OpenAI released a policy proposal calling for the United States to lead the formulation of global technical standards for frontier AI, incorporating RSI into the discussion. It proposed the need to measure the degree of automated research within AI companies, define which automated research processes should trigger human review, and establish unified incident classification and reporting thresholds; the proposal also clarified that these standards do not constitute a license, nor are they mandatory pre-release reviews.
Meanwhile, Dario Amodei, co-founder and CEO of Anthropic, published "We Must Pace the Frontier" on September 12, 2026, advocating for proactively slowing down the pace of model capability improvements to buy time for alignment, interpretability, and safety verification. The "International AI Safety Report" released in February 2026, authored by over a hundred independent experts, listed "loss of control" as a distinct risk category, while noting that experts are highly divided on its likelihood and that current systems do not yet possess the capability to trigger such risks. However, not everyone agrees with this sense of urgency.
Some researchers point out that the current empirical evidence for so-called RSI remains weak, and most achievements are more accurately termed "constrained automated optimization" rather than true recursive self-improvement. Packaging constrained optimization as RSI will not only mislead public perception but may also catalyze unnecessary regulatory overreactions. The essence of this debate is a clash between two different senses of time.
One side believes the capability curve is steepening, and the window for verification and governance is narrowing; the other side argues that verification itself is the bottleneck, and it is premature to discuss loss of control before we can reliably prove that "the next generation is indeed stronger." Regardless of which side one stands on, one thing is certain: as AI begins to participate in designing its own next generation, the question humanity needs to answer is no longer "what can it do," but "can we still understand what it is doing."
- End -