Google DeepMind, together with Singapore’s AI Safety Institute, OpenMined, AVERI, and MLCommons, piloted what it describes as the first double-blind evaluation framework for proprietary AI models. The setup runs inside Google Cloud’s Confidential Computing (Confidential Space), a cryptographically verified secure enclave, so evaluators’ test prompts stay hidden from the model’s developers while the model’s weights stay hidden from the evaluators. DeepMind says the approach is designed to prevent benchmark contamination and strengthen trust in AI capability assessments.