Google DeepMind pilots double-blind evaluation to curb benchmark contamination

Google DeepMind says it has piloted what it calls the world's first double-blind evaluation of a proprietary, frontier-class AI model. For the pilot, DeepMind tested a Gemini Flash Lite model against confidential benchmarks, working with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. The test ran inside Confidential Space, part of Google Cloud's Confidential Computing portfolio, which DeepMind used to cryptographically verify that the evaluators' benchmark questions and Google's own model weights each stayed private to their owner.
The pilot targets what DeepMind calls benchmark contamination: if a model has already seen a benchmark's test questions before being scored on them, the resulting numbers cannot be fully trusted, since a high score might reflect a model repeating material it has already seen rather than showing real capability. As AI models grow more capable, DeepMind argues, policymakers, researchers, and enterprises increasingly need assurance that a reported benchmark score reflects a model's true capabilities and safety, and that assurance breaks down if a model provider can see evaluation questions in advance and tune against them.
Before this pilot, DeepMind says, evaluating a proprietary model from outside forced a tradeoff: either the evaluator handed its test prompts to the model's maker, risking that the maker could see the questions in advance, or the model's maker handed its model weights to the evaluator, risking exposure of its intellectual property. The double-blind design is meant to remove that tradeoff: within the cryptographically secured environment, the evaluator cannot see Gemini's model weights, and Google cannot see the evaluator's test prompts, so each side's material stays confidential to its owner while the evaluation still runs.
DeepMind says zero-logging protocols and contractual safeguards have already kept external test prompts confidential for some time; what is new is layering technical, cryptographic guarantees on top of those protections, which it calls a major step forward for secure model evaluation. It flags cybersecurity testing and evaluations run by government bodies as cases where this guarantee matters most. The post discloses no numeric results, scores, or outcomes from the Gemini Flash Lite test, and it does not name the specific confidential benchmarks used; it describes the evaluation method, not the findings. DeepMind calls the effort a pilot and says it hopes the approach establishes a new frontier for model oversight that helps the wider industry build safer, more reliable, and more widely trusted AI systems, stopping short of claiming that outcome as already achieved.
Key facts
- Google DeepMind piloted what it calls the world's first double-blind evaluation of a proprietary, frontier-class AI model.
- The pilot tested a Gemini Flash Lite model against confidential benchmarks inside Confidential Space, part of Google Cloud's Confidential Computing portfolio.
- External partners named: the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
- The cryptographic design keeps both sides blind: the evaluator cannot see Gemini's model weights, and Google cannot see the evaluator's test prompts.
- DeepMind discloses no numeric results, scores, or the specific benchmarks used; the post describes the method, not the outcome, and frames the effort as a pilot.
Why it matters
Benchmark scores are how the industry backs up claims that a model is capable or safe, but a model that has already seen the test questions can inflate its own score without demonstrating real capability, a failure mode DeepMind calls benchmark contamination. External evaluators had relied on zero-logging protocols and contractual promises to keep test prompts confidential; DeepMind now adds a cryptographic guarantee on top of that, verified through Google Cloud's Confidential Space rather than taken on trust. That shift matters because it turns evaluation integrity from a policy commitment into something a third party can technically verify, which DeepMind positions as necessary for policymakers, researchers, and enterprises to trust that a reported score reflects a model's true capability and safety.
Who it affects
The pilot concerns frontier AI labs whose proprietary models undergo outside evaluation, and the independent organizations that run those evaluations, here the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. DeepMind singles out cybersecurity testing and evaluations run by government bodies as the cases where this guarantee matters most, since those evaluations tend to involve especially sensitive prompts on the evaluator's side and especially sensitive model weights on the developer's side. More broadly, it affects anyone who relies on a benchmark claim, including policymakers and enterprises weighing whether a model's stated capabilities and safety properties can be trusted.
How to use it
This is not a product, tool, or API a reader can adopt directly. It is an internal pilot DeepMind ran with the four named partner organizations, using Google Cloud's Confidential Space. DeepMind does not describe how another lab or evaluator could start a similar double-blind evaluation, what it costs, or whether the method will be offered more broadly beyond this pilot. Readers wanting the technical mechanics are pointed to a separate technical report that DeepMind says covers its methodology and findings, though the post itself does not summarize what that report contains.
How solid is it
The account comes entirely from DeepMind's own blog post, written in the company's voice, with no statement or reaction quoted from any of the four partner organizations and no named individual attached to the work. The 'world's first' claim is DeepMind's own characterization, not confirmed by an outside party. The post discloses no numeric results, scores, or outcomes from the Gemini Flash Lite test, and it does not name the confidential benchmarks used, so there is no way to independently assess whether the double-blind method changed anything about the outcome compared with prior evaluation approaches. DeepMind itself frames the result as a hope rather than a demonstrated achievement, writing that it hopes the pilot establishes a new frontier for model oversight, and the post carries no date beyond 'today.'
Risks and caveats
Because no evaluation results are published, it is not possible to tell whether double-blind testing changed the benchmark scores, caught contamination that would otherwise have gone undetected, or simply reproduced what earlier evaluation methods already found. Whether cryptographically secured evaluation becomes standard practice across the industry, or remains a Google-run pilot with these four named partners, is left open, and so is who would bear the infrastructure cost of running routine evaluations inside a confidential-computing environment going forward. None of the four partner organizations is quoted describing its own experience of the pilot, so the account of how the process actually worked comes from Google alone, and readers cannot check it against the evaluators' side.
“The evaluator cannot see the Gemini model weights, and Google cannot see the evaluator’s test prompts.”
— Google DeepMind