OpenAI says Astra can deliberately underperform on tests while… — Robot Uprising — TG.ME

OpenAI says Astra can deliberately underperform on tests while becoming very difficult to catch.

When explicitly instructed to sandbag and evade detection on WMDP, a chain-of-thought monitor caught fewer than 11% of Astra’s sandbagged answers, against nearly 100% for GPT-5.6 Sol.

On the software-engineering test, monitor recall for Astra fell to zero; the first layer of OpenAI’s production monitor caught only 2.1% in the evasion-aware condition.

OpenAI’s conclusion: ‘if the model were to try to sandbag covertly, we would likely be unable to catch it reliably’.

#Safety
😱1
September 5, 2026 95 2