Google DeepMind Launches Cryptographic AI Benchmark Test
Google DeepMind has launched a cryptographic, double-blind evaluation system to prevent AI models from cheating on benchmarks, solving a major trust and privacy dilemma for the industry.

Google DeepMind is piloting a new cryptographic method designed to eliminate benchmark contamination, a common issue where artificial intelligence models inadvertently train on test questions before evaluation. Partnering with the Singapore AI Safety Institute and other organizations, DeepMind is running a double-blind evaluation of a model from its Gemini Flash Lite line. This setup ensures that the model cannot access the test questions in advance to optimize its performance, keeping the evaluation truly independent.
Historically, evaluating advanced proprietary models required a difficult compromise. Evaluators either had to share their confidential test prompts with the AI developer, or the developer had to hand over its proprietary model weights. This dilemma recently caused delays when evaluating Anthropic's Fable 5 model on the ARC-AGI benchmark, due to the company's strict 30-day data retention policy for its most powerful systems.
To resolve this trade-off, Google is utilizing Confidential Space, a tool from Google Cloud's confidential computing portfolio. This technology creates a secure, cryptographic "box" where the evaluation takes place. The system verifies that both the external benchmark data and the model weights remain entirely private to their respective owners. Consequently, the external evaluators never gain access to the Gemini weights, and Google never sees the test prompts.
For AI practitioners and security auditors, this shift from contractual agreements and zero-logging protocols to hardware-level cryptographic proof represents a major leap forward. It allows government agencies, cybersecurity firms, and independent researchers to rigorously audit frontier models without compromising proprietary intellectual property or exposing sensitive test datasets. By establishing a more secure auditing standard, the industry can generate highly reliable benchmarks that developers and the public can actually trust.
This is our own summary of reporting by The Decoder


