Overview of Recursive Self-Improving (RSI)
RSI refers to the procedure where an AI system analyzes, updates, and redesigns itself to enhance next-generation’s capability and efficiency. It creates a feedback loop where each generation performs better than previous generation. But personally speaking, RSI is … probably another hype.
Table of Contents
1. RSI
1.1. What is RSI
In short, RSI simply refers to the process where AI improves itself through analysis, update, and possibly redesign in various aspects, like model architectures, systems, kernels.
1.2. Levels of RSI
Depending on the extent of automation, RSI can be roughly divided into 5 levels, from lowest to highest.
- AI only executes what human designs
- AI chooses the strategy how to improve
- AI plans for what to learn and experience in the next iteration
- Environment interaction and self-adaptation
- Recursive generation
1.3. Basic Components
- Base Model
- Target to optimize
- Improver, Meta Agent. Responsible for analysis
- Sandbox, Evaluator
- Retained System State
- Experience Store, Replay World.
1.4. Enterprise Practices
However the concrete solution of RSI is, all practices fit into the pattern of “Perform, Evaluate, Modify, Redeploy”.
1.5. Challenges
- Reward hacking causes “fake improvement”
- Failure to monitor CoT, causing “fake reasoning”
- Human-in-the-loop becomes the bottleneck
1.6. Future Direction
- Sample Efficiency
- Persistent Memory
- Evaluator Co-Evolution