Цирк какой то When Hugging Face's incident response team tried to… — whargarbl — TG.ME

Цирк какой то
When Hugging Face's incident response team tried to analyze the attack, they fed real exploit payloads and command-and-control artifacts into frontier commercial models (Claude, GPT-class) to help reconstruct the timeline.

The safety guardrails blocked them.

The models couldn't distinguish an incident responder analyzing a breach from an attacker constructing one. Real exploit code in the prompt triggered the same refusal whether you're attacking or defending. The team had to pivot to GLM 5.2 running on their own private infrastructure, which also kept the stolen credentials from leaving the environment.

Read that sequence again. The attacker's agent had no guardrails. It ran unrestricted, autonomous, at machine speed. The defenders' AI tools were restricted by the same safety layers designed to prevent exactly the thing the attacker had already done.

The attacker moved at machine speed. The defenders were slowed by their own safety rails. That asymmetry is the entire lesson.
Reddit
From the better_claw community on Reddit
Explore this post and more from the better_claw community
👀2
July 21, 2026 380 4