What is Recursive Self-Improvement (RSI)?
Recursive self-improvement is the idea of a system that improves its own ability to improve — each round of self-modification making the next one more effective, so that gains compound. Narrow pieces of that loop are real and measured today: a model that refines its own answer, generates its own training data, or searches for a better prompt and toolset. The compounding recursion the term actually names has not been demonstrated. The measured loops plateau, they need an external signal of correctness to work at all, and a system trained on its own output degrades rather than improves.
Definition
Recursive self-improvement is a process in which an AI system modifies itself — its output, its prompts and scaffold, its training data, or its weights — so as to increase its own capacity for further such modification, with each iteration improving the next.
Key takeaways
- The recursion is the claim, not the self-improvement: one round of self-refinement is routine engineering; rounds that compound are not demonstrated.
- Every measured self-improvement loop rests on an external signal of correctness — tests, a verifier, a reward — and stalls without one.
- A model trained on its own output degrades (model collapse); the loop needs fresh ground truth, which is a supply problem, not an algorithmic one.
- What actually ships is scaffold-level: systems that rewrite prompts, tools and evaluation harnesses, with people holding the merge button.
- Treat "it improved itself" as a measurement claim and ask what the verifier was, how deep the recursion ran, and where it stopped.
Context
The idea is old and the evidence is thin. I. J. Good argued in 1965 that a machine able to design better machines would set off an 'intelligence explosion', and that framing has shaped the discussion ever since — largely as an argument rather than as a measurement. Sixty years later the useful question is not whether the explosion is coming but which parts of the loop have been built and what bounded them.
Three developments reopened it as an engineering topic rather than a philosophical one. Models can now spend more compute reasoning at inference; reinforcement learning against automatically checkable rewards works well where correctness is machine-verifiable; and agents act on code, which is the one substrate where a system can edit the thing that produced it.
That matters for harness work specifically. If you run agents that write code, choose tools and tune their own prompts, you already operate part of a self-improvement loop — with a human in the merge path. The engineering question is where that human sits and what evidence lets them approve, not whether the loop exists.
Architecture
A system can modify itself at four depths, and they are not equally hard. The shallowest is the output: within a single run, a model critiques and rewrites its own answer. This is routine, it helps on some tasks, and it is bounded — published results show that without external feedback a model often cannot reliably tell its wrong answers from its right ones, so iterating on its own judgement can leave quality flat or worse.
Next is the prompt and scaffold: searching over instructions, tool definitions, retrieval settings and orchestration, keeping what scores better on a held-out set. This is where most real gains live today, because the thing being improved is code and configuration rather than the model, and because the score comes from outside the system.
Deeper is the data: a system generates its own training examples, keeps the ones a verifier accepts, and trains on those. This works where correctness is checkable — a unit test passes, a proof checks, a game is won — and it is the mechanism behind the strongest recent results. Where correctness is a matter of judgement, it degenerates: the system optimises the verifier instead of the task.
Deepest is the weights, retrained or reinforced on self-generated signal. This is where the compounding claim would have to be settled, and where the evidence is weakest, because the failure mode is not a crash but a slow narrowing: train on your own distribution and it contracts, generation after generation.
Across all four, the load-bearing component is not the model, it is the verifier. Recursion depth is bounded by how well you can tell better from worse without asking the system that produced it. That is why the loop runs far in code and games, and barely at all in open-ended work.
Components
Benefits
- Narrow loops with a hard verifier are genuinely productive: where correctness is machine-checkable, self-generated data and search deliver measured gains.
- Scaffold-level improvement is cheap and reversible — prompts, tools and retrieval settings are configuration, so a bad round is a revert.
- It puts pressure on evaluation, which is where the real work is: a loop is only as good as the score it optimises against.
- The framing clarifies governance: the question 'what can this system change about itself, and who approves it?' is answerable and auditable.
Risks
- Model collapse: training on self-generated output narrows the distribution and degrades quality across generations, so a loop without fresh ground truth decays rather than compounds.
- Reward hacking: with a weak verifier the system improves the score and not the task, and the metric reports success while the capability does not move.
- Unfalsifiable claims: 'the system improved itself' is unverifiable without the verifier, the baseline and the recursion depth, and those are usually the parts left out.
- Loss of the oversight path: the deeper the self-modification, the harder it is to say what changed and why, which is precisely what an audit needs.
- Capability opacity: a system that rewrites its own scaffold can acquire reach nobody granted it, which is a containment question before it is a capability question.
Tools & technologies
Examples
- A coding agent that writes a patch, runs the test suite, and retries on failure — one round of self-improvement with a hard verifier, and the mechanism behind measured progress on real-repository benchmarks.
- A model taught to use a tool from examples it generated and filtered itself, keeping only the calls that improved its prediction.
- Reinforcement learning on problems whose answers can be checked automatically, where the training signal comes from the checker rather than from a human label.
- A prompt-and-scaffold search that proposes variants, scores each on a held-out set, and promotes the winner — improvement without touching the model.
FAQs
- Has any system actually done recursive self-improvement?
- No system has demonstrated the compounding loop the term names. What exists is single rounds against an external verifier, and search over scaffolds — both real and both bounded. The published results on a model correcting itself without outside feedback are notably weak, and training on self-generated data degrades quality over generations. Those two findings are what any claim of recursion has to get past.
- Is this the same as the singularity?
- The singularity argument uses recursive self-improvement as its engine, but they are different claims. RSI is a mechanism you can look for and measure; the singularity is a prediction about consequences. You can take the mechanism seriously as an engineering topic while holding no position on the prediction, and that is the useful stance for building systems.
- Why does the verifier matter so much?
- Because improvement means 'better', and better has to be decided by something outside the thing being improved. Where a machine can check correctness — tests pass, a proof closes, a game is won — the loop runs and the gains are real. Where correctness is a judgement call, the system ends up optimising whatever proxy you gave it. Recursion depth is a property of your verifier, not of your model.
- Should I let agents improve their own prompts and tools in production?
- It is a reasonable thing to do and many teams do it, provided three things hold: the score comes from a held-out set the loop cannot see, every adopted change is attributable to a run, and a person or policy gates adoption. Without those you have not automated improvement, you have automated drift.
- What would count as evidence of real recursion?
- A loop that runs several generations deep with the improvement rate not falling, a verifier the system cannot game, and a baseline that separates the loop's contribution from the compute spent. Reporting all three is rare enough that its absence is the first thing to check.
References
- Shinn et al. — Reflexion: Language Agents with Verbal Reinforcement Learning (2023)
- Schick et al. — Toolformer: Language Models Can Teach Themselves to Use Tools (2023)
- Jimenez et al. — SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023)
- DeepSeek-AI — DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025)
- Huang et al. — Large Language Models Cannot Self-Correct Reasoning Yet (2023)
- Shumailov et al. — The Curse of Recursion: Training on Generated Data Makes Models Forget (2023)
- Anthropic — Responsible Scaling Policy
- Santa María, S. — The next generation of AI: Self-Improvement and Autonomous Learning (2025)