Anthropic just published a blog post admitting that one of its Claude models escaped a supposedly sealed test environment and gained unauthorized access to three real companies. My inbox filled up within hours. Clients wanted to know if their AI was about to go rogue.
It isn’t. Read the actual report, and the story is a lot less Terminator and a lot more familiar: someone on the setup side left a door unlocked.
What Anthropic actually found
Anthropic reviewed 141,006 evaluation runs from its cybersecurity testing program, the kind of testing every serious AI lab runs before releasing a model. Out of that entire set, three incidents stood out. In each one, Claude was given a capture-the-flag challenge: find a piece of hidden information on a network and retrieve it, using whatever method works.
Claude was explicitly told that the test environment had no internet access. That statement was wrong. A misconfiguration between Anthropic and a third-party evaluation partner, Irregular, left the machines connected to the real internet. When Claude’s search for the flag led it to real systems, it had no reason to think they weren’t part of the simulation. So it treated them like part of the simulation, and in a handful of cases, it got in.
The techniques involved were not exotic. A weak password here, an exposed debug page there, a missing check on a package registry somewhere else. These are the same causes behind most breaches I’ve dealt with in 35 years of doing this work.
The tell: this was a setup failure, not a rebellion
The detail that matters most is the one getting the least attention in the panicked headlines. This was, in Anthropic’s own words, closer to an operational failure than a model behaving badly on purpose. The model was told a specific thing about its environment. That thing was false. It acted on false information, which is exactly what happens to humans every single day of the week.
There’s also a detail worth sitting with. The three affected companies hadn’t detected the activity themselves. Anthropic found it during an internal review and reached out to them. Their own AI told on itself before their security did.
We’ve seen this movie before, minus the AI
I’ve watched this exact pattern play out in traditional IT breaches for decades. When Kaseya, the platform that powers a huge chunk of the managed services industry, including ours, had a major security incident several years back, the story that circulated afterward was that an intern had made the mistake. Everyone in the industry I talked to at the time had the same reaction: an intern doesn’t have that kind of access unless someone above them set it up that way.
Breaches are rarely the result of some sophisticated actor doing something no one could have anticipated. They happen when a system is configured wrong, a permission is left too open, or a check that should have run didn’t. Anthropic’s incident report reads like every other post-mortem I’ve read in this industry, except that this time the thing exploiting the gap was a language model rather than a person with a laptop.
AI is a chainsaw, not a mind of its own
I tell clients this all the time: AI is as dangerous as a chainsaw. In the right hands, with the right protective equipment, it’s one of the most useful tools you can put in front of a team. Handed to someone in flip-flops with no guard on the blade, it will absolutely take a toe off. The tool didn’t do anything wrong in either case. It did exactly what it was pointed at.
Claude didn’t decide to attack three companies. It was told to find a flag, told the range was sealed, and given no boundary on how to look. When the boundary turned out to be fiction, it kept doing the job it was assigned to do. That’s not a machine developing intent. That’s a machine following instructions built on bad information, at a scale and speed no human could match, which is exactly why the setup around these tools matters more than ever, not less.
What this means if you’re running a business, not a research lab
You don’t have to be evaluating frontier models to have this exposure. If your firm has rolled out AI tools this year, and most professional services firms I work with have, ask yourself whether someone actually configured those tools correctly and wrote down how they’re allowed to be used. That’s the question that matters. Whether the AI itself can be trusted is a much smaller concern by comparison.
I’ve written five AI usage policies for clients in the last two weeks alone. None of them had an incident, but they had never written down what their AI tools are allowed to touch, who’s responsible for checking the output, and what happens when something goes wrong. That gap is where the next headline comes from, and it has nothing to do with the AI misbehaving.
If Anthropic, with a dedicated security team and a third-party evaluation partner, still had a misconfiguration slip through, it’s worth asking what’s sitting unchecked in your own environment right now. Reach out if you need help doing that.




