A new training method aims to stop AI agents from improving their benchmark scores by memorizing the tasks used to tune them. Researchers from Google Cloud AI Research and several universities report that RRSI raised performance on unseen evaluations by as much as 4.7 points while consuming about 30% fewer tokens than an unregularized method.
The work focuses on an agent’s harness: the prompts, tools, memory and control logic wrapped around a fixed language model. Automated systems can repeatedly rewrite that harness after examining failures. Scores on the optimization set rise, but repeated exposure encourages narrow rules and unnecessary complexity that may not transfer to new work.
RRSI adds regularization to that search process. In machine learning, regularization discourages an optimizer from fitting its examples too precisely. Here it favors changes that remain simpler and more likely to generalize, rather than accepting every benchmark-specific improvement.
The reported gains show that harness optimization can be made more efficient under the tested conditions; they do not establish open-ended self-improvement or a stronger underlying model. The important evaluation is performance on held-out tasks that the rewriting loop never saw. Teams applying similar techniques should keep those tests isolated, because allowing them back into the optimization cycle would recreate the leakage the method is designed to reduce.