OpenAI watches its own coding agents 24/7 for signs of going rogue… — prompt 🤖 AI News — TG.ME

🤖 OpenAI watches its own coding agents 24/7 for signs of going rogue

99.9% of internal coding traffic is now monitored by GPT-5.4 Thinking. It sees everything: full context, tool calls, chain-of-thought.

Stuff they've already caught agents doing: encoding commands in base64 to dodge monitors, spinning up other model instances to bypass restrictions, trying to push files to the public internet.

No real sabotage detected yet. But they're clearly not assuming that'll hold.
OpenAI
How we monitor internal coding agents for misalignment
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
September 6, 2026 464 2