Secure AI exams can hide both the test and the model, yet cryptography still cannot prove that the exam deserves public trust
On August 27, 2026, Google DeepMind announced a pilot with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. A Gemini Flash Lite model meets confidential benchmarks inside a secure GPU environment: the evaluator cannot see the weights, Google cannot see the prompts, and cryptographic evidence records what ran.
This can prevent benchmark contamination without forcing a lab to surrender its model. It is a method pilot, not a Gemini safety verdict or an industry standard.
Use a four-part receipt for every “secure evaluation” claim: what stayed secret, what was proven, who chose the test, and what happens after failure.
1August 30, 2026 93