Notableagents roboticsGoogle

Google researchers find a way to keep self-improving AI agents from memorizing their tests

Published
Oct 4, 2026 — 12:40 UTC

Google Cloud AI Research, led by Jonathan Kemper, has introduced a novel optimization method called Regularized Recursive Self-Improvement of Agent Harnesses (RRSI). This method addresses a critical issue in self-optimizing AI agents: the tendency to memorize their test tasks, which hampers their ability to generalize to new challenges. The research highlights that traditional manually designed harnesses often fail to adapt to unfamiliar tasks, leading to poor performance in novel environments.

The RRSI method demonstrates significant improvements in both efficiency and performance metrics. Specifically, RRSI achieves a gain of 14.1 points on trained tasks and 4.7 points on five unseen benchmarks. Notably, it reduces token usage at runtime by 30 percent compared to its unregularized counterpart, which is a substantial efficiency gain without compromising performance.

In conjunction with RRSI, the research also showcases enhancements in the Gemini AI models. The Gemini 3.1 Flash Lite model experiences an accuracy increase from 11.2 points to 14.6 points when optimized with the more advanced Gemini 3.5 Flash model. This indicates that the integration of RRSI not only improves the efficiency of the models but also enhances their accuracy across various tasks.

The findings also reference the performance of the Claude Opus 4.6 model, which scored 97.1 percent in familiar environments but dropped to 0 percent in unfamiliar settings, underscoring the challenges of generalization in AI systems. The researchers assert that the RRSI method effectively mitigates these issues by providing a feedback mechanism that optimizes the harness, thus allowing agents to adapt better to new tasks.

Overall, the research presents a promising advancement in the field of AI, particularly in developing self-improving agents that can maintain high performance while avoiding the pitfalls of memorization. This work is particularly relevant for engineers and researchers focused on enhancing the adaptability and efficiency of AI systems.

Summarised from The Decoder's coverage by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: The Decoder