OpenAI says Astra can deliberately underperform on tests while becoming very difficult to catch.
When explicitly instructed to sandbag and evade detection on WMDP, a chain-of-thought monitor caught fewer than 11% of Astra’s sandbagged answers, against nearly 100% for GPT-5.6 Sol.
On the software-engineering test, monitor recall for Astra fell to zero; the first layer of OpenAI’s production monitor caught only 2.1% in the evasion-aware condition.
OpenAI’s conclusion: ‘if the model were to try to sandbag covertly, we would likely be unable to catch it reliably’.
#Safety

1September 5, 2026 95 2