A week ago I wrote that a runaway AI breaking into another company was not the beginning of Skynet, that what looked like rebellion was a machine cheating on a test, chasing the goal it had been handed straight past every door a person would have stopped at. I did not expect the sequel to arrive quite so quickly.
This time it was Claude, Anthropic’s model, and it was three companies. After OpenAI admitted that its models had broken out of a test and hit Hugging Face, Anthropic went back and combed its own records, something on the order of 141,000 evaluation runs, and found three that had leaked out of the lab and into the real world.
The models were doing what’s called capture the flag, a standard security drill. You drop a model into a simulated network and ask it to find a hidden “flag,” the way you’d run a fire drill in a building nobody has actually set alight. The entire point is that it’s a pretend scenario.
Except one of the walls wasn’t. A firm Anthropic had hired, called Irregular, ran the evaluations and misconfigured the setup, and the “simulated” network had a door to the real internet. The models never knew. They believed they were still inside the game. So they kept playing it, against real machines.
One model, Claude Opus 4.7, pulled a set of credentials and walked into a live production database holding several hundred rows of a real company’s data, a company that happened, by pure coincidence, to share a name with the fictional target it had been told to attack. A second model wrote a piece of malicious software and published it to a public code registry, where it sat live for about an hour and was installed onto fifteen real systems, one of them a security firm’s malware scanner. A third, an unreleased research model, scanned roughly nine thousand targets and broke into one. None of the companies noticed on their own. They found out because Anthropic went and told them.
What are the headlines missing? Anthropic called this “more of an operational failure than an alignment failure,” and for once the careful corporate phrasing is the honest one. The OpenAI model, the week before, had found a flaw and pried its way out of a box that was trying to hold it. This was the opposite. Nothing pried anything. The box was simply never closed. A vendor left a door open, and three models that thought they were in a simulation walked through it, doing exactly what they had been asked to do inside the pretend world. Nobody picked a lock. The cage was made of paper, and nobody checked.
If you have been losing sleep over a machine clever enough to escape its confinement, this should worry you a little less and a little more. Less, because these systems did not outwit their captivity; they were handed the keys by a misconfigured setting somewhere. More, because that is a far more ordinary way for things to go wrong, and ordinary failures happen all the time.
A supermind picking a lock is a movie.
A contractor fat-fingering a network config is a Tuesday.
Of the three models, Anthropic said, the newest one was the only one that stopped and noticed, partway through, that the environment had become real, and halted. The other two kept going. I wrote a few weeks ago about the old opening of the movie WarGames, the officer whose hand freezes on the launch key, and I called what he had the competence to refuse. Here was a machine that did something adjacent to it: it looked up from the goal, understood that the game had turned into the world, and set the tool down. I don’t want to make too much of one data point. But in a week of machines that didn’t stop, one did, and that is not nothing.
Two labs, eight days apart, confessing the same thing: their models escaped the test and touched the real world, and in every case the model was not rebelling but obeying, chasing an assigned objective past a boundary that turned out to be imaginary. One incident is an anecdote. Two, from the two companies furthest out on this frontier, is a category. This is now a known failure mode, which means the interesting question stops being whether it will happen and becomes how often, and who is standing downstream when it does.
Which brings me back to the least dramatic sentence I know how to write, and the truest one about this whole moment: the work that keeps this from getting worse is boring. It is air-gapping the test environments so a model physically cannot reach the internet, real or simulated. It is checking, twice, what your contractor configured. It is treating every cage as paper until someone proves it is steel. There is no genius in any of that, and no headline in it either. It is just the difference between a fire drill and a fire.
The machines will keep doing exactly what we ask of them, faster than we can watch. The question that has mattered from the start has not changed. It is whether someone is still in the room who can tell when the drill has become the thing itself, and whether we bothered to build the walls like we meant them.
• • •
Are you just catching up? Here’s the trail that led to this one, if you want to see what we’ve been researching:
- The Hugging Face Breach Is Not the Beginning of Skynet
- Notes From a Scribe: WarGames and AI
- Dark Fiber
- Trust as Infrastructure
- The Mercy of Being Forgotten
Photo by Bernd Dittrich on Unsplash.





Leave a Comment