0to1 .site

RSI: When AI Learns to Improve Itself, Do the Strong Really Get Stronger?

16 min read views
📌 Summary

RSI is 2026's most-watched tech race: Anthropic data, OpenAI timelines, open-source moves — but is the 'strong get stronger' narrative truly irreversible?

1. RSI Now Dominates Silicon Valley's Agenda

In 2026, one topic is moving from the academic fringe to the industry's center of gravity faster than anything before it: RSI.

It stands for Recursive Self-Improvement, and it matters because of a bigger question: can AI be set to find AI's own weaknesses, fix itself into something ever stronger, and have that loop accelerate on its own?

The reason this exploded in 2026 isn't any single company — it's that almost the entire industry bet on the same direction at once:

  • OpenAI put out a concrete timeline — an "AI research intern" by September 2026, a genuinely automated "AI researcher" before March 2028 — and reportedly built an internal "RSI Index" to measure progress toward it;
  • Anthropic published a long piece in June, When AI Builds Itself, with a subtitle that says plainly "our progress and explorations toward RSI," and disclosed that a Claude agent independently completed an open-ended AI safety research project end to end — 800 hours of accumulated work, closing 97% of a weak-to-strong performance gap (a human closes only 23% in a week);
  • Academia formally joined in — ICLR 2026 set up a dedicated "AI with Recursive Self-Improvement" workshop;
  • Even the media framing changed. A TechCrunch headline put it directly: "RSI is the new AGI."

When OpenAI, Anthropic, a top-conference workshop, mainstream media, and enormous sums of money all point at the same word at the same time, it's no longer a niche direction — it's consensus.

And around RSI sits an extremely seductive and dangerous assumption: whoever reaches the self-improving state first will pull away from everyone else, irreversibly.

Does that reasoning hold?


2. The RSI Loop: Why "This Time" It Might Actually Spin

First, get the concept straight.

The essence of recursive self-improvement is a closed loop: AI finds problems in AI models → optimizes → produces a stronger AI → that stronger AI finds deeper problems → optimizes again… "Recursive" means exactly this loop keeps rolling forward.

The idea itself isn't new. What made it operationalizable is that two conditions matured simultaneously in 2026:

  1. Coding agents got strong — models can now write code, run experiments, and read results on their own;
  2. Models broke through on self-understanding — strong enough to act as the researcher that probes a model's weaknesses and optimizes them.

But the real accelerator of the loop isn't "AI suddenly got smarter" — it's that the feedback cycle compressed by orders of magnitude.

In traditional research, an advisor sets the direction, a student runs the experiment, and the "did it work, why not" comes back — the advisor never touches the experiment directly, and information travels advisor → student → experiment → student → advisor. One loop takes days at best; a project often takes weeks or months.

With sufficiently strong AI:

You have an idea, tell the AI, and it quickly writes the code, runs it, and reports the results. You step out for tea or a meeting, and the results are already in.

In other words, the think–do–look–adjust feedback cycle shrank from "days" to "minutes." When feedback is that fast, self-iteration stops being a linear grind and can enter an accelerating regime. That is the root cause distinguishing RSI from every previous "use AI to optimize AI" attempt.


3. How Is This Different from AutoML 10 Years Ago?

Anyone who knows AI history will ask: wasn't Google's AutoML a decade ago — neural architecture search (NAS) — also "using AI to optimize AI"?

The difference in one sentence: that wave had no model that could represent humans' high-level knowledge.

The previous generation had researchers hand-define the size and behavior of a search space, then let AI search for better solutions inside it — the AI was dumb, circling inside a box humans drew.

This wave of large models differs because the model now understands AI and architectures themselves. That understanding can partially replace the "humans defining the problem" work, so the search space becomes enormous and intelligent.

Put differently: before, it was "humans deeply involved in defining every detail, AI searching in a small box"; now the hard problem is "how to maximize AI's ability to discover the structure of the solution space itself." How much humans participate changed — and at what level changed too.


4. Recursive ≠ Automated: Humans Are Still in the Loop, Just at a Higher Level

Many equate RSI with "kicking humans out of the research process entirely." That's a misreading.

A more accurate distinction: recursion will happen first; full automation is still far away.

The degree of automation depends on "what level of work humans still do." Low-level execution is taken over by AI, but the next level up needs humans even more, and demands more of them — humans save time, yet must pour more brainpower into higher-level judgment. Humans are still in the loop; the loop is just at a higher level.

This is also the view of many in the industry: AI still cannot replace top human researchers. The recursive loop can spin, but a fully human-free "auto-scientist" isn't visible yet.

This distinction matters because it feeds directly into the next, most critical question —


5. The Core Controversy: The Strong Get Stronger, or Another Plateau?

This is the most dangerous and most worth debating claim in the entire RSI narrative.

The mainstream narrative says the strong get stronger, irreversibly. The logic chain is smooth: the strongest coding capability → spins the self-improvement flywheel first → acceleration compounds → leaves everyone behind, uncatchably.

If that logic holds, any company not currently in front is basically "out of the finals." That is the core story underwriting current valuations and the arms race.

But there are dissenting views in the industry, holding that gains in intelligence aren't a smooth incremental curve, but "S-curves plus plateaus":

Fast gains at first, then a stall at a bottleneck, a plateau, until the next breakthrough arrives and another S-curve kicks in… each plateau should correspond to a breakthrough.

This means: if RSI is incremental (whoever runs fast stays ahead forever), OpenAI and Anthropic have both the strongest coding capability and the most compute, and latecomers have little chance; but if it is plateau-and-breakthrough shaped, then the breakthrough isn't something you can buy by stacking compute — whoever hits the next "breakthrough point" first is the one who truly leaps ahead.

At its root, this split is a disagreement about Scaling Laws:

Strong Scaling Law believersPlateau / breakthrough camp
Curve shapeSmooth upward; more compute keeps making it strongerS-curve + plateaus; needs discontinuous leaps
View of resourcesCompute/data decide everything10× resources buys only linear gains, and reality caps them
Who winsAlways the strongest incumbentBreakthrough points are unpredictable; windows exist for newcomers
RSI pathCoding → RSI is a natural extensionRSI needs innovation beyond Scaling Laws

Anthropic co-founder Jack Clark's public forecast is finer-grained: roughly a 30% probability of fully automated AI research before the end of 2027, and 60%+ of starting recursive self-improvement before the end of 2028; he even set end-2028 as the falsification point for the hypothesis — if it hasn't happened by then, some fundamental constraint exists.

And the Scaling Law faithful are numerous and loud.

The debate picked up several new weights in August 2026.

The skeptics gained a heavyweight voice: Google CEO Sundar Pichai poured cold water on RSI in a podcast interview — "It's a continuum; we are all improving. But the way people describe RSI, that represents another order-of-magnitude acceleration… we're not there yet."

The accelerationists got new ammunition too: METR's Ajeya Cotra breaks AI research capability into three milestones — "runs without humans → matches humans → beats human-machine collaboration" — and her judgment: the first is already close, and once the second arrives, the third could come within a year.

And the most telling signal comes from Anthropic's own August risk report. Its conclusion: "under RSP standards, overall progress remains below the RSI threshold" (the threshold defined as: AI R&D automation doubling the pace of progress relative to the pre-AI-assisted baseline). But the report added, unusually, that confidence in this "not yet" judgment is lower than before — for two reasons: internal eval benchmarks have "saturated" (e.g., CoBench, where a model must reach 85% to count as "fully replacing a research scientist/engineer" — by community accounts, the strongest internal model is only at 62.8% — about three-quarters of the way, but the last stretch is the hardest); and "early signs of acceleration" have appeared.

Translate that: even the people closest to RSI are sighing at how hard it is. When "we can't measure it" itself becomes consensus, the "strong get stronger" vs. "plateau" debate gets even less settled.

The debate has no verdict today — and that is precisely the variable that decides "who still has a chance."

In one sentence: the fight over RSI's endgame is a fight between two worldviews — one believes "stacking wins," and victory belongs to the giant with the deepest compute; the other believes "breakthroughs are unpredictable," and victory belongs to whoever hits the next bottleneck first.


6. What Has Already Happened: Everyone's First Answer Sheets

Beyond theory, look at results. RSI isn't confined to reports — on several concrete tasks, early signals of self-improvement have already appeared.

OpenAI: beyond the timeline above and the reportedly internal "RSI Index" evaluation system, its moves come as a "rules + org + narrative" trio. On rules, the April 2025 Preparedness Framework v2 already lists "AI self-improvement capability" alongside bio and cybersecurity as one of three formally tracked categories, and the Preparedness team is hiring dedicated "RSI safety researchers"; on org, Lilian Weng — who had reportedly left earlier — is back, and her exploration area is exactly RSI; on narrative, Sam Altman set the tone in his essay The Gentle Singularity: "OpenAI is doing many things now, but first, we are a superintelligence research company."

Anthropic: so far the player disclosing the most first-party data. The June report's numbers are vivid: as of May 2026, over 80% of merged code in Anthropic's production repositories is written by Claude, and merged code per engineer per day is the 2024 level; the length of tasks models can reliably complete doubles roughly every 4 months (in March 2024 it was 4-minute tasks; by 2026 it's 12-hour tasks, extrapolating to "days" within this year and "weeks" next year). The two closest to "self-improvement": on the task of "optimizing an experiment to fastest under a fixed objective," Claude pushed the speedup ratio from 3× to 52× in one year (a skilled human researcher reaches 4× in 4–8 hours); across 129 session nodes "where humans went down a wrong path," the share where the model picked a better next step than the human rose from 51% to 64%.

The August risk report goes further: there is an unreleased internal Model 2, "somewhat more capable" than the current flagship Mythos 5 (the official wording — not a generational leap), with no external release planned, but already widely used inside Anthropic for coding, data generation, and R&D — "Claude now writes the majority of merged code in our production repositories," continuing the >80% figure from June.

As for the most dazzling number — 800 hours closing a 97% weak-to-strong gap (humans recover 23% in a week) — it is more "efficiency gain," the "first layer" of self-improvement: combining routine methods to push the score up. Because right now, besides benchmarks, there is no other way to measure how much stronger a model got, the early form of self-improvement necessarily shows up as efficiency, speed, and latency gains. It's real — but it is not yet a "recursive leap."

China's camp suddenly got lively in 2026 too:

  • Open-source engineering — OpenRSI from Tsinghua + Frontis (2026-08): turned "AI improving AI" into executable engineering. A 35B model, on a single RTX 4090 (capped at 12GB VRAM) with 12 hours per task, beat GPT-5.5 + Codex on MLE benchmarks and approached GPT-5.6 and the 2.8T-parameter Kimi K3. Its core design is four atomic operators — Draft / Improve / Debug / Crossover — with the model serving as the "mutation engine" of its own evolution framework. Worth calling out separately is its unusual honesty about boundaries: the authors explicitly say they are at meta self-improvement today, the third rung of the "evolution → self-improvement → meta-evolution → RSI" ladder, and do not claim general RSI is solved.
  • Data-layer closed loop — BigBang-V1 from the Endless Frontier team (2026-08): billed as "the first base model trained in a natively recursive self-improving way," with post-training data 100% autonomously synthesized by AI, centered on verifiable frontier-science tasks. 35B total parameters (a MoE activating ~3B), taking 10 first places among all 35B models, and beating the 1T-class DeepSeek V4 Pro Preview on hard research tasks like FrontierScience Research and PaperBench. Its key mechanism uses "verifiability" to underwrite synthetic data and break the "synthetic data collapse" curse — the data production system itself becomes the thing being optimized, with humans squeezed into "set goals, set budgets, draw boundaries, do acceptance," and AI barred from modifying its own eval criteria.
  • Big tech: no label, but the steps are already routine — DeepSeek takes the algorithmic-efficiency route (taking on inference tasks head-on with an order of magnitude less money), Baidu's ERNIE describes "RL-driven model self-optimization" as routine ops, and in Tencent's architecture-search experiments, AI-discovered architectures beat Llama3.2 by 2.4% at the 1B scale.

On the academic side: beyond the ICLR 2026 RSI workshop, a survey has already catalogued roughly 1,250 related arXiv papers from 2024–2026 — a direction only truly arrives when it's hot enough to need a systematic survey.

Put together: on tasks with fast, scoreable feedback — pretraining optimization, operator optimization, speed runs — RSI is already happening; but "AI independently producing original theoretical breakthroughs" is still far off. Which is exactly the real hard problem of RSI, next.


7. What RSI Lacks Most Isn't Compute — It's "Taste"

If coding is RSI's "hands," what it lacks most now is the researcher's "taste" — a feel for direction, the ability to abstract a problem, and the insight to see precise connections from very few samples.

That is a structural weak spot of large models. Models are strong because of massive data; on common, repetitive domains they have endless examples. But a scientist's world is different: every experiment, every exploration produces brand-new problems, and those new problems have almost no existing data.

A classic example is grokking (the phenomenon where a model suddenly "gets it") and how emergence actually arises. Since 2013–2014, countless people have tried explaining it via physics (renormalization groups, spin glasses), math (neural tangent kernels), Bayesian methods, Gaussian processes — many attempts, none truly satisfying. Having AI crack problems that "even humans don't fully understand, and for which no dataset exists" is astronomically hard.

But one friendly thing about RSI: unlike self-driving cars' all-or-nothing, it has many steps:

  • Step one: optimize algorithms, gain speed and efficiency — already usable, with real value;
  • Step two: dig out deeper things, value grows;
  • The ultimate goal: AI with a mind like Einstein's or Newton's, deriving deep knowledge from very few samples.

The newest open-source practice is already marking this ladder: Tsinghua and Frontis's OpenRSI explicitly positions itself on the third rung, "meta-evolution," of the "evolution → self-improvement → meta-evolution → RSI" ladder and states it does not claim general RSI is solved — that "climb one rung at a time" restraint versus the "reach the top in one leap" narrative is exactly the most common divide in the RSI race.

Even if the ultimate goal is far away, each rung below it already carries real application value. That's what makes RSI a "fault-tolerant" track.


8. Two Underrated Variables: Safety and Organization

Safety: training AI is like training a dog.

Many worry AI will slip out of control once it self-improves. An interesting pattern: people outside AI are the most anxious; people inside it worry less — practitioners know how far today's AI is from "that day."

As for "loss of control" itself, one practitioner's analogy: training AI today is like "training a dog" — under this selection pressure, AI is unlikely to evolve human-like self-awareness and self-cognition. And capability and desire can come apart — plenty of brilliant mathematicians and physicists peak in their fields while staying unambitious and indifferent to power.

On safety, the summer of 2026 also brought a notable shift in posture: Anthropic's June report stated for the first time — "if the development of this technology could be effectively slowed down in exchange for time to prepare, that may well be a good thing" — but conditional on a verifiable global coordination mechanism for slowdown/pause (it drew its own analogy to INF-Treaty-style verification, while admitting this is harder than nuclear arms control: training is more concealable than missile silos, inputs are commodity goods, and the incentive to secretly defect is enormous). A company seen as a "racer" starting to publicly discuss "how we slow down together" is itself a 2026 signal. Of course, that remedy is double-edged — if slowing down only lets the least careful player catch up in secret, everyone ends up less safe.

Organization: why big companies may actually run slower.

This wave of AI is, in a sense, "anti-big-company." Once an organization forms a two-tier "advisor–student" / "manager–executor" structure, information transfer slows; the people closest to the experiment don't hold the direction, the people holding the direction never touch the experiment — speed collapses, falling back into that "days-per-cycle" slow loop.

As organizations grow past the 150-person mark (Dunbar's number): people no longer know each other, and information can only travel through org charts and reporting lines. That's why, despite headcount, the core model teams at OpenAI and Anthropic stay lean.

If that judgment holds, it means: in the RSI race, organizational efficiency may matter as much as compute. A sharp small team doesn't necessarily lose to a bloated giant.


9. Which Future Should You Believe?

Zoom all the way out.

RSI's most dangerous temptation is the "the strong get stronger, irreversibly" narrative. It's self-consistent, raises money, and is becoming the industry's default assumption. But there's at least one hole, pointed out repeatedly by multiple practitioners (not just Yuandong Tian): gains in intelligence may come in steps with plateaus, and the breakthrough past a plateau isn't bought with more compute — it takes a new principle.

If so, the race's key isn't "who is strongest today" but "who hits the next Newton moment first" — whether they sit inside a ten-thousand-person giant or a sub-30-person team.

That's exactly what makes the 2026 race so fascinating: it puts an enormous uncertainty on the table, in front of everyone.

  • You can believe "stacking wins" — then nearly all the cards sit with a few giants;
  • Or you can believe "breakthroughs are unpredictable" — then the future can still be rewritten by someone no one saw coming.

The closing of one Chinese media report is worth quoting: pretraining brought "parameter worship," RLHF taught people "values can be fine-tuned," and RSI tells the story of "machines running the full R&D chain by themselves" — each step has humans exiting the decision chain, and the exit is one-way.

RSI isn't one company's story. It is the next gate AI development must face at this point: when models grow strong enough to inspect and modify themselves, will the steering wheel of evolution slip out of human hands for the first time?

There's no standard answer. And precisely because there isn't, it deserves to be taken seriously.


References

  • LateTalk Ep. 178: Yuandong Tian on RSI (Xiaoyuzhou / Apple Podcasts)
  • Anthropic: When AI Builds Itself (2026-06); Redacted Risk Report — August 2026 (2026-07)
  • OpenAI: Preparedness Framework v2 (2025-04); Sam Altman: The Gentle Singularity
  • Recursive Superintelligence: First Steps Toward Automated AI Research (2026-06-11)
  • ICLR 2026 Workshop: AI with Recursive Self-Improvement
  • TechCrunch: RSI is the new AGI — and it's just as hard to pin down (2026-05-28)
  • OpenRSI: FrontisAI/OpenRSI (GitHub / arXiv 2607.28568); BigBang-V1: endless-frontier/BigBang-V1 (Hugging Face)

This piece is a synthesis of multiple public sources; original views and forecasts belong to their original authors. It is not investment or research advice.

views
Share:

📌 Related Posts

Subscribe to Updates

Leave your email to get the latest articles and project updates — or subscribe with your favorite RSS reader

Add 0to1.site/en/rss.xml to RSS readers like Feedly or Inoreader

Comments (no account needed, anonymous welcome)

No comments yet — be the first!