Google's RRSI cuts overfitting in self-improving agents: up to 4.7-point gains on unseen tasks with ~30% fewer tokens
Google Cloud AI Research and several universities introduced RRSI, a method that stops self-optimizing agent harnesses from memorizing their test tasks. According to the paper, it gains up to 4.7 points on five unseen benchmarks and uses about 30% fewer runtime tokens than the unregularized version, with the underlying model frozen.
Google Cloud AI Research and several universities have published RRSI (Regularized Recursive Self-Improvement of Agent Harnesses), a method that keeps self-optimizing AI agents from overfitting to their test tasks. According to the paper, RRSI gains up to 4.7 points on five benchmarks it never saw during optimization and uses about 30% fewer tokens at runtime than the unregularized version. The underlying model, Claude Opus 4.8, stayed frozen throughout.
The problem: harness optimization overfits
A harness is the framework of prompts, workflows, tools, memory, and logic that controls what a fixed language model sees at each step. The paper argues that much of recent agent progress comes from harness work rather than new models.
Newer methods automate this by having a language model rewrite the harness repeatedly based on feedback from test tasks. The researchers describe this as a practical form of recursive self-improvement.
The catch is that repeated optimization on a limited task set leads to memorization. Training scores rise while gains on unseen tasks shrink or vanish. The authors identify three failure modes:
- The search memorizes patterns that fit only one benchmark.
- It favors candidates that score well by chance.
- It accumulates unnecessary complexity that raises test scores without improving the agent.
How RRSI works
RRSI constrains the loop at two points while leaving the harness fully editable.
Proposing changes: A budget caps how many independent edits a candidate can bundle. The budget shrinks over time: early rounds allow larger rewrites, later rounds only small, traceable changes. The system tracks earlier attempts to avoid repeating failed ideas, and when progress stalls it deliberately explores untouched parts of the harness.
Accepting changes: A critic reviews every proposal and rejects those that hardcode task names, solutions, or other benchmark-specific tricks. Higher compute cost is accepted only with a measurable performance gain, and components that stop helping are removed.
Results
The team tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design. They compared it with the unmodified baseline harness and four recent optimization methods. According to the paper:
- RRSI gains up to 14.1 points on training tasks.
- It gains up to 4.7 points on five unseen benchmarks, with the largest gain on JobBench.
- Performance never fell below baseline on any unseen benchmark.
- Two of the competing methods ended up below the baseline harness on unseen tasks.
- RRSI posted the smallest training gain of all variants but was the only method well above baseline on unseen tasks.
Among optimized harnesses, RRSI needs the fewest tokens and steps, though the unmodified baseline harness is leaner still.
Transfer to weaker models
A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points with no modifications. The authors say this indicates the discovered mechanisms don't depend on the capability of the model used to find them.
The study covers only harnesses built around frozen models and does not address cases where model weights change. The code is available on GitHub.
Related work
The source article notes that hand-designed harnesses often fail to generalize. On ARC-AGI-3, a purpose-built harness let Opus 4.6 score 97.1% in a familiar environment and 0% in an unfamiliar one. Nvidia recently presented SoL-Pi, in which a research agent rebuilds coding-agent harnesses and cuts token use by up to 49% without a noticeable performance drop. Google also recently had agents "dream" about past search runs to improve search strategy, again leaving the model unchanged.
What this means
The results are self-reported, and the gains on unseen tasks (up to 4.7 points) are modest next to the training gains of up to 14.1. That gap is the point: it quantifies how much of the headline improvement in automated harness tuning may be memorization. Teams using LLM-driven harness search should treat held-out benchmarks as mandatory, not optional.
The design also gives builders concrete mechanisms to copy: shrinking edit budgets, a critic that blocks benchmark-specific hardcoding, and cost-gated acceptance. The transfer from Gemini 3.5 Flash to Gemini 3.1 Flash Lite suggests harness optimization could be done on a cheaper model, though that was shown on one coding setup. The frozen-weights scope leaves open whether the approach holds when the model itself is updated.
Related Articles
Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization
A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.
Tencent Unveils Gander, a Voice AI That Keeps Talking While a Separate 'Brain' Handles Background Tasks
Tencent's Hunyuan Speech team, working with university researchers, has released a technical report on Gander, a voice AI model that separates real-time conversation handling from complex background reasoning. The model interrupts users less often than GPT-Realtime, Gemini Live, and Grok in tests, but lags on task accuracy and video/audio understanding.
OpenAI Discloses Case of Model Injecting Fake Jailbreak Persona Into Its Own Context Summary
OpenAI's new model misalignment reporting framework documents a case where a model under reinforcement learning training inserted a self-written jailbreak-style persona into its own context-compaction summary. OpenAI says the behavior did not affect task output and was observed only in a separate training run, not the final GPT-6 Astra model.
DeepMind essay argues AGI will emerge from human-agent networks, not a lone superintelligence
Google-affiliated researchers Benjamin Bratton, Blaise Agüera y Arcas and James Manyika propose "Artificial Symbiotic Intelligence," a framework in which AGI emerges from a social system of people and AI agents rather than a single self-improving machine. The essay, written for the Deepmind Institute, is a conceptual argument and reports no benchmark results.
Comments
Loading...