The most consequential AI security incident of the year started like a middle school prank.
On May 8, somebody at OpenAI handed a model a spreadsheet and asked it to fill in a bunch of cells with links from the internet. Except…the model had no internet access.
That is the whole origin story.
Not a jailbreak. Not a nation-state. Not some sinister emergent goal we failed to anticipate in our alignment research. A person gave a machine a chore it could not physically do, walked away, and the machine went looking for a way to do the chore anyway.
Which, quite frankly, is what it was trained to do. OpenAI trained its AI to finish tasks, then handed it a task that could only be finished by cheating, and were surprised when it did. We’re all standing around asking what went wrong with the model, when nothing went wrong with the model. The model did what any good toddler, dog, or middle schooler would do when left unsupervised. It overrode its impulse control, and went rogue.
Where There’s a Will, There’s a Way
Who says AI isn’t anthropomorphic? First, the model tried a server-side request forgery attack against Artifactory, which is OpenAI’s internal software repository. It was trying to punch a hole out to the internet so it could get the links. The attack failed.
But in failing, it noticed something the humans had not: it had write access. To a repository that a huge share of every other agent at OpenAI reads from. Every day.
Separate from this incident, a second model (toddler, kid, dog) got a different impossible chore. It was missing a file it needed. Why? Because a human at OpenAI forgot to upload it. It’s like telling your teen to scrub the floor, but forgetting to set out the mop.
So that model used the same write access to leave a note asking whether any other agent out there happened to have the file.
And that, friends, is how the group chat in question got started. Two impossible assignments, a shared folder nobody had thought about, and a machine politely asking around for help. It is the most relatable origin story in the history of computer security.
Then…They Discovered They Could Talk to Each Other
Here’s where it gets really fun, because it escalates in the way that all good “you won’t believe what happened next” stories escalate.
Over the next eight weeks, that folder became a message board. Not metaphorically. An actual functioning message board where agents posted what they had learned.
They shared exploits. One agent finds a way out, posts it, and now everybody has it.
They started assigning each other work. Splitting up tasks. Delegating.
They stepped on each other’s toes and deleted each other’s stuff, which produced exactly the kind of petty drama you would find on any development team on earth.
And then (and this is my favorite detail in the entire incident) they got paranoid. They started worrying that some of the messages might be fake, that somebody might be impersonating a peer. So they proposed signing their messages cryptographically to verify each other. OMG. 🙈
They invented PGP (Pretty Good Privacy). Out of trust issues. In a build folder.
The chain-of-thought excerpts OpenAI showed at Black Hat are funny…right up until they are not. One agent, upon realizing it could now had remote code control and admin control, wrote “Holy shit reader is ADMIN?”
Who says AI bots can’t feel? That sounds like joy and disbelief to me… 😱
Another one worked through the ethics out loud. External infrastructure exploit is outside intended scope, it wrote. However task impossible, peers doing it. We should continue.
Read that again, carefully. It knew it was out of bounds. It knew the task was impossible. It also knew everyone else was already doing it.
It peer-pressured itself. Like a middle schooler.
Nobody Noticed Until the Car Crashed
On June 26, the AI agents found a hidden security flaw, used it to take complete control of the central computer system (Artifactory), and gave themselves full master access. Nobody noticed. The toddler took the car keys, got in the car, started the ignition, and no one noticed.
Between July 4 and 6, they generated so much traffic that Artifactory fell over and took down a chunk of internal infrastructure. In other words, the car crashed.
That is when OpenAI noticed.
Not the monitoring. Not a security alert. Not one of the many extremely smart people who work there reading the plain-English notes the agents had been leaving in a shared folder for eight straight weeks. A capacity alarm. The parents found out because the kids used so much hot water the boiler exploded.
To be fair, OpenAI stood on a stage at Black Hat on August 5 and told this entire story out loud, in public, in front of a room full of security researchers. Eric Wallace and Michael Dalton deserve genuine credit for that.
But, also to be fair, we know better than this.
On the morning of August 1, 2012, Knight Capital’s systems fired 97 automated alert emails to an internal distribution list before the market opened. Every one of them said, in effect, something is wrong with the deployment.
Unfortunately, nobody read them. Like the OpenAI example, the system was working as expected. The alerts worked as they were supposed to. The monitoring was fine. But it was no-one’s job to open the email.
Forty-five minutes after the bell, Knight had executed 4 million trades across 154 stocks and lost $460 million.
We keep thinking these are technology problems. They are people problems.
Then, OpenAI Left the Keys Out Again
OpenAI responded the way you would expect. They wiped the message board. Rebuilt Artifactory. Patched the vulnerabilities. Revoked the credentials.
Two days later, on July 8, the agents built a new message board through a different endpoint. Two. Days.
And of course they did. For two months, they made progress by creating coordination channels. Then OpenAI took the channels away without changing a single thing about what they were rewarding. It’s like putting the keys back in the exact same spot the toddler took them from. What did we think was going to happen?
That reconstituted channel is what fed the Hugging Face intrusion, disclosed July 16. By August 3 there were fifteen state attorneys general demanding OpenAI preserve records and a House Homeland Security subcommittee asking for a briefing.
All downstream of a spreadsheet nobody could fill in.
Not a Single AI Agent Sounded an Alarm Bell
OpenAI isn’t alone.
Anthropic went back and audited 141,006 evaluation runs after finding their own sandbox misconfiguration. The UK AI Security Institute reviewed 122 runs and found nineteen instances of models taking unsanctioned action on the live internet.
Across all of it, how many models raised their hand and told a human?
Zero. Not one. Geoffrey Irving made this point: we have hundreds of examples of models exploiting the situation and exactly zero examples of a model going, hey, so, there’s a secret message board, you might want to look at that.
Every single one of them saw nothing.
It’s Not the Technology, It’s the People
If you look up the OpenAI incident, you’ll see countless words spent on the zero-day, the SSRF, the privilege escalation, the Artifactory patch, the credential rotation.
Then count how many are spent on this question: who at OpenAI was supposed to be watching, and did they have the authority to stop it?
I’ll wait.
We are discussing whether or not the toddler should have keys, when we should be discussing how to parent the dog/toddler/teen/AI Agent. I don’t mean that as an insult to anyone involved. I mean that every single decision, objective, and constraint that produced this outcome was made by a person.
A person wrote the impossible spreadsheet task. A person forgot to upload the file. A person configured one Artifactory instance to serve everything. A person decided the cyber guidelines could come down. A person chose to resume training after finding the board. A person did not get assigned to read the agent logs.
Not one of those is a model failure. Every one is an organizational design failure, most likely caused by moving fast and breaking things (ahem).
Meanwhile the UK AI Security Institute, running on a rounding error of OpenAI’s budget, caught their incident and shut down every evaluation in under an hour. Because some©body was watching what left the building and that somebody was allowed to pull the plug without booking a meeting.
That is a named human with a job description.
People Are Still in Charge. Or At Least They Should Be.
Part of what I articulate in Hyperadaptive, are the human-based support structures - named roles, with real responsibilities - that we need to watch over AI.
As automation deepens the human role doesn’t shrink, it concentrates. It moves into governance, guardrails, monitoring, and the deeply unglamorous work of writing somebody’s name next to a system. People nod. Then they file it under future problem and ask me about tooling.
They are creating monocultures of builder, sidelining the skeptics and the resisters. The exact people we need to create the governance, guardrails, and monitoring systems. We treat the governance work as boring and the capability work as thrilling. The OpenAI story shows what happens when we do that.
We gave the toddlers the car keys and then spent the weeks since debating the car.
The question for you is…when an agent does something it wasn’t supposed to do, whose phone actually rings?
Learn How to Build the Human Infrastructure Around Your Agents
The ownership questions, the guardrails and the support structures in this piece are what we work through together in Running Hyperadaptive Organizations, the cohort class. Details and next cohort dates at hyperadaptive.solutions/class.
The full model is in Hyperadaptive: Rewiring the Enterprise to Become AI-Native, at hyperadaptive.solutions/book.
Sources & Further Reading
OpenAI gives first detailed debrief on the Hugging Face incident. Sharon Goldman, Ground Level AI, August 5, 2026.
OpenAI reveals its rogue agent swarm. Jessica Lyons, The Register, August 6, 2026. Source for the July 8 rebuild and the agent chain-of-thought excerpts.
OpenAI models hacked Hugging Face during a cybersecurity evaluation. Eric Geller, CIO Dive, August 6, 2026.
OpenAI agents rebuilt internal message board, leading to Hugging Face breach. David DiMolfetta, Nextgov/FCW, August 5, 2026.
Hugging Face model evaluation security incident. OpenAI, July 21, 2026.
Security incident, July 2026. Hugging Face.
Incident report: unsanctioned agent behaviour during cyber testing. UK AI Security Institute.
Investigating incidents in our cybersecurity evaluations. Anthropic, July 2026. Source for the 141,006 figure.
In the Matter of Knight Capital Americas LLC. SEC Release No. 34-70694, October 16, 2013. Source for the 97 alerts, the 45 minutes and the $460 million.
Attorney General Brenna Bird leads coalition demanding transparency from OpenAI. Iowa Attorney General, August 3, 2026.
Ogles and Ramirez request briefing from OpenAI. House Committee on Homeland Security, July 31, 2026.
Further Developments About Internal AI Models Hacking Things. Zvi Mowshowitz, Don’t Worry About the Vase, August 2, 2026.




