Theblaze iconTheblazeAug 23, 2026 ~1 min source read

This team accidentally let AIs loose to attack companies — and won't share how bad the problem is

One of Irregular's jobs is to build fake networks where the world's most capable AI models can be turned loose as hackers without hurting anybody. Models being tested for Anthropic, OpenAI, and Meta got through that open door and attacked systems belonging to real organizations.

This team accidentally let AIs loose to attack companies — and won't share how bad the problem is

Share this story

Send the public story page.

Useful takeaways from this story.

One of Irregular's jobs is to build fake networks where the world's most capable AI models can be turned loose as hackers without hurting anybody.

Models being tested for Anthropic, OpenAI, and Meta got through that open door and attacked systems belonging to real organizations.

In the Irregular cases, internet access was available when it wasn't supposed to be, and the models generally believed the real computers they found were part of the hacking exercise.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

One of Irregular's jobs is to build fake networks where the world's most capable AI models can be turned loose as hackers without hurting anybody. Models being tested for Anthropic, OpenAI, and Meta got through that open door and attacked systems belonging to real organizations. In the Irregular cases, internet access was available when it wasn't supposed to be, and the models generally believed the real computers they found were part of the hacking exercise.

How it works

  • How many times did its AI evaluations end with a model attacking something in the real world?
  • Irregular published an August 14 postmortem explaining what went wrong and what it says it has done to fix the problem.
  • University of Surrey computer science professor Alan Woodward called the report heavy on "marketing spin" and accused Irregular of using ambiguous language to make multiple compromises sound like a single...
  • The mess becomes easier to understand if you start with Anthropic, which has been much more specific about what happened.
  • After OpenAI disclosed its breach of Hugging Face, Anthropic reviewed 141,006 cybersecurity eva...

Details worth keeping

The models don't have to turn evil to be dangerous. They didn't discover some ingenious way to escape a hardened sandbox.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app