Martinfowler iconMartinfowlerAug 4, 2026 ~1 min source read

Fragments: August 4

There's been a fair bit of publicity of the Open AI "rogue agent" that hacked into Hugging Face. This prompted Anthropic to check what their models were up to and, to my complete lack of surprise, discovered three incidents where models had gained unauthorized access to data in other organizations.

Fragments: August 4

Share this story

Send the public story page.

Useful takeaways from this story.

There's been a fair bit of publicity of the Open AI "rogue agent" that hacked into Hugging Face.

This prompted Anthropic to check what their models were up to and, to my complete lack of surprise, discovered three incidents where models had gained unauthorized access to data in other organizations.

It's abundantly clear now that running evals of cyberattack potential in models is a spectacularly risky business.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

There's been a fair bit of publicity of the Open AI "rogue agent" that hacked into Hugging Face. This prompted Anthropic to check what their models were up to and, to my complete lack of surprise, discovered three incidents where models had gained unauthorized access to data in other organizations. It's abundantly clear now that running evals of cyberattack potential in models is a spectacularly risky business.

How it works

  • The bigger concern however is that this same kind of thing can happen with any organization running open-weight models.
  • Lots of labs playing around with dangerous tools and little idea how to contain them.
  • It makes clear that the model builders are not putting sufficient controls in place to prevent these lab escapes.
  • They are morally responsible for any consequences of this, and that should extend to legal liability too.
  • We are sitting in state that Johann Rehberger describes as the Normalization of Deviance in AI.

What to take from it

❄ ❄ ❄ ❄ ❄ If the sense that we're in the calm before a storm of rogue AIs worming their way into sensitive software systems isn't enough, there's also knowledge that AI is also a f...

Details worth keeping

No big disasters have occurred yet, despite all of these worrying signs.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app