The Information Machine
Following·since 27 Aug 2026·Day 2·7 sources·updated 28 Aug 2026

Google DeepMind's double-blind AI evaluation pilot

The gist

Google DeepMind Pilots First Double-Blind Frontier AI Evaluation

Benchmark contamination is a recognized trust problem for AI evaluations, and double-blind setups remove one of the key reasons to distrust published scores. The organizers argued the results show it is technically feasible for policymakers to require deep, secure third-party access to AI models.

The full picture

Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons completed a pilot double-blind evaluation of Gemini 2.5 Flash-Lite using Google Cloud's Confidential Space. The evaluation is the first of its kind for a proprietary frontier model: hardware attestation verifies exactly what code runs before either party's assets enter the enclave, so the evaluator cannot see Gemini model weights and Google cannot see the evaluator's test prompts. AVERI encrypted evaluation prompts with a private key invisible to all other parties, jointly ran the evaluation inside a Google-configured enclave using OpenMined software, and alone decrypted and graded outputs.

The pilot targets benchmark contamination: when a model trains on the same questions used to grade it, its scores measure memorization rather than capability. Previously, high-stakes external evaluations required a tradeoff between exposing test prompts to model developers or exposing model weights to evaluators; the double-blind setup eliminates that compromise. A 2024 pilot by OpenMined, Anthropic, and the AI Security Institute previously demonstrated the secure enclave mechanism using GPT-2 and a five-row evaluation; this pilot advances to a production model and a more comprehensive evaluation.

How it developed
28 August 2026

Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons announced August 27 the completion of a pilot evaluation of Gemini 2.5 Flash-Lite using a double-blind setup inside Google Cloud's Confidential Space, where AVERI encrypted test prompts invisible to Google while Google kept model weights invisible to AVERI.

The setup targets benchmark contamination, in which a model trained on evaluation questions produces scores reflecting memorization rather than capability. The organizers said policymakers considering audit requirements under California SB 315 and the EU General-Purpose AI Code of Practice should require deep, secure third-party access to AI models.

27 August 2026

Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons published results of the first double-blind evaluation of a proprietary frontier model, Gemini 2.5 Flash-Lite

Sources
2 more sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free