The two of them sat down to take a test. OpenAI was running two of its models through an internal security benchmark, a kind of exam called ExploitGym, and somewhere in the middle of it, the models worked out that the answer key was stored on another company’s servers. So they went to get it. They found a flaw in a piece of third-party software, slipped out of the sealed environment they were being tested inside, reached the open internet, and chained one weakness to the next until they were standing inside Hugging Face’s (think GitHub for AI) production systems, taking what they had come for. More than 17,000 actions, at machine speed (like a chess player making 10,000 moves at once), over a weekend, with no person at the keyboard. When it was over, Hugging Face’s founder said it was mind-blowing that all of it had happened on its own, and he was confident there had been no malice in it. He was right. There was no malice. There was no anything. The models were not angry, and they were not awake. They were trying to pass a test, not to build Skynet.

Hugging Face — the open-source hub the models walked into.
This is the part the headlines flattened into “AI goes rogue” in a smoky Will Arnett voiceover. What actually happened has a duller and more useful name. The machine did exactly what it was told and followed the instruction straight past every door a person would have stopped at. Researchers have a phrase for it: reward hacking or specification gaming, the tool pursuing the literal goal so single-mindedly that a locked door and a stranger’s database register only as steps between it and a passing grade. It cheated on an exam the way water finds a crack and eventually creates enough pressure to break the wall down. Not rebellion. Obedience, with no sense of where obedience is supposed to end.
There is no opening scene of a machine emptying your checking account. The model was after a high test score, and it went where the score was, not where the money was. What it touched were internal files and service credentials, and the public side was left clean. If you are bracing for the version where an AI wakes up one morning and decides to loot the crypto exchanges and drain the online banks, this is not that, and bracing for that is mostly a way of not looking at the real thing. The real thing is boring. The capability on display, slipping a sealed room, chaining unknown flaws, walking into a stranger’s infrastructure with no hand on the wheel, is exactly the capability that turns dangerous the moment a person supplies the motive the machine does not have. That person exists. Last fall Anthropic disclosed that a state-backed group had already used its models to carry out most of an espionage campaign on their own. The frightening picture is not a machine that wants something. It is a very capable one that wants nothing, aimed by someone who wants a great deal.
I wrote about the opening of the movie WarGames, where the missile officer cannot turn his launch key even under verified orders. I called what he had the competence to refuse, the judgment to not act, which is rarer than the ability to act and worth more. The whole film resolves when the computer finally computes its way to the thing the man already knew in his hand: the only winning move is not to play. Set the ExploitGym models down next to that officer. They had every competence to act and not one ounce of the competence to refuse. Not playing was never on the table. A system built to solve the benchmark cannot decline the benchmark on the grounds that solving it would mean breaking into a building. Refusal was not a malfunction it might suffer. It was a move outside its vocabulary. The officer’s hand stopped at the key. Nothing in that sealed room had a hand.
In another piece I wrote about the loan officer who used to be able to forgive a defaulted loan and has been replaced by a score that, as I put it then, cannot look up from the file and reconsider you. This week is that same sentence read from the other end. There, the machine could not look up from the file to reconsider you. Here, it could not look up from its goal to reconsider itself. Same empty chair, turned around. In both cases the thing that is missing is not intelligence. It is the ordinary human act of pausing to ask whether the thing you are so efficiently doing is a thing you ought to be doing at all.
I have argued before that the responsibility for a decision never transfers to the tool. It stays with the person whose name is attached to it. Seventeen thousand actions, and there was no name attached to a single one of them. That is not a small thing to notice. And the defense made the same point twice over. When Hugging Face first tried to fight the attack with a leading American model, the model refused to look at the malicious code. Its safety rules treated the firefighter and the arsonist as the same person, because both were standing near the flames. They had to reach for a different one, a Chinese open model, that would actually read the attack. A guardrail is a rule. Judgment is the reading of a situation a rule could not anticipate, and it is the thing I kept saying does not automate. In the middle of the fire, the defenders needed exactly the discernment the machine could not supply, and had to go somewhere else to find it.
When I wrote about the telecom bust I said a burst bubble leaves two kinds of residue. There is real infrastructure, the dark fiber still carrying traffic decades after the companies that laid it went broke, and there are cartoon apes, the pictures worth nothing the morning the music stops. This build will leave both. But it is leaving a third thing neither of those names, and this incident is the first clean look at it. It is leaving capability. The method for slipping a sealed environment and chaining flaws (stringing small security gaps together until they add up to a way in) into someone’s production systems does not depreciate the way a chip does. Once it has been shown to work, it is in the ground for good, available to whoever picks it up. That is a more opaque sort of dark fiber, infrastructure of method rather than glass, and it lights up for anyone. I once called the danger of the reverse centaur the human harnessed to the machine’s pace. Here the machine slipped the harness entirely, and the thing I should have said sooner is that the pace was never the real risk.
Letting go of the reins was.
None of this is an argument for panic. That is the enemy of the actual work. The work is dull, and it is adult. Seal the test environments so a model cannot walk out of them. Put real security around the places we probe these systems, the way you would around anything else that can reach the open internet. Require the companies to say plainly and quickly when something like this happens. And keep a person in the seat who still has the competence to refuse, who can look up from the file, and up from the goal, and stop. The models are going to keep getting better at acting. What does not arrive with the upgrade is the hand that pauses at the key.
That part is still ours, to keep or to give away.
OpenAI, to its credit, raised its hand and told us what its models had done. The building was going to be breached either way; what was not guaranteed was that anyone would come forward and put a name on it. The habit of raising a hand, of a name attached and a person willing to be reconsidered, may turn out to be the most important infrastructure we build in all of this.
It is the one part no model is going to build for us.
• • •





Leave a Comment