🤖 GPT-6 Astra still cracks under hidden attacks
OpenAI claims GPT-6 Astra blocks 99.99% of direct prompt injections.
Bury the attack inside a document instead, and it fails 8.5% of the time.
Claude Opus 5 beats it here, failing only 4.8% of scenarios.
Security theater matters less than what happens when agents read the open web unsupervised.
Autonomous agents are being shipped faster than the defenses that should gate them.

September 4, 2026 43 1