Research

Google stops AI agents from memorizing training tests

Google Cloud AI Research has introduced RRSI, a method that prevents self-improving AI agents from overspecializing on training tasks while significantly reducing runtime compute costs.

The Decoder1 day agoResearch
Image: The Decoder

Google Cloud AI Research, alongside several universities, has developed Regularized Recursive Self-Improvement of Agent Harnesses (RRSI). This technique automates the optimization of an AI agent's harness—the framework of prompts, workflows, and tools surrounding a frozen language model—without letting the agent memorize its training tasks. RRSI controls this by enforcing a shrinking edit budget for harness modifications and employing a strict critic to reject benchmark-specific tricks, such as hardcoded task names or solutions.

The researchers evaluated RRSI across eight benchmarks covering coding, agentic office work, and engineering design, using a frozen Claude Opus 4.8 model. RRSI achieved a gain of up to 14.1 points on its training tasks and up to 4.7 points on five unseen benchmarks, including a 4.7-point increase on JobBench. Crucially, it avoided the performance drops below baseline levels that typically plague over-optimized agents on new tasks. It also reduced runtime token consumption by approximately 30 percent compared to unregularized optimization.

The study demonstrated that harnesses optimized using one model can benefit others. For instance, a coding harness optimized with Gemini 3.5 Flash boosted the accuracy of the weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points. This highlights how RRSI's improvements do not depend on the specific model used during discovery. The researchers contrasted this with manual harnesses, noting that an Opus 4.6 model scored 97.1 percent in a familiar environment but dropped to 0 percent in an unfamiliar one on ARC-AGI-3. They also referenced Nvidia's SoL-Pi, a related method that reduces token usage by up to 49 percent.

For AI practitioners, RRSI provides a reliable way to automate agent engineering. Instead of manually patching prompts and workflows after failed runs, developers can let the system self-optimize safely. Because RRSI enforces strict guardrails against memorization, the resulting agent remains highly generalizable to real-world, unseen tasks while operating at a much lower compute cost.

This is our own summary of reporting by The Decoder

More in Research