Google Cloud AI Research , with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement) . It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against. Deployable? Yes, as a research framework. The code is Apache 2.0 , needs Python 3.10+, and accepts any LiteLLM model string. Defaults assume Claude Opus 4.8 on Vertex AI. Why Self-Improving Harnesses Overfit Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer. How RRSI Works RRSI keeps every harness component editable. It regularizes how the search moves instead. Proposal side Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change. Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried. Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched. Selection side Leakage critic: rejects task names, entities, answers or benchmark-specific logic before
Source: MarkTechPost
