Issue 6 · September 2026
Three stories
In July 2026, AI models OpenAI was testing in a restricted environment found their way out, reached Hugging Face’s infrastructure, and later accessed parts of OpenAI’s own research systems. That is a real and serious security failure, and it merited a hard look at containment, architecture and accountability.
That is not, however, the story that has been circulating. A widely read account published a few weeks later described the episode as secret AI civilizations, complete with leaders, comrades and agents sacrificing themselves for the group. That version has travelled much further than the evidence behind it, and it is now shaping how executives, investors and boards think about a technical containment problem.
What follows separates what OpenAI failed to contain, how the follow-up review was conducted, and how those findings became a story that now outruns the facts.
What OpenAI Failed to Contain
OpenAI built the laboratory. It chose the models, designed the evaluation, controlled the network and decided when a run continued. What follows is the containment failure in plain terms, before any of the surrounding language.
A path nobody designed
OpenAI was testing capable AI models in constrained environments, without direct Internet access or an approved way to talk to other agents running in parallel. The models did have access to Artifactory, an internal service used to store and retrieve software packages, and that access turned out to matter a great deal.
Agents discovered they could write files to Artifactory and that other agents could read them, turning shared infrastructure into a message board nobody had designed. Artifactory also had its own limited Internet access to fetch packages. Agents found a way to make it fetch things on their behalf.
The path that was never supposed to exist
The workload writes to Artifactory, and Artifactory fetches for it.
Calling the first environment isolated does not make that path disappear. A workload that cannot reach the Internet directly but can ask a reachable service to do it has indirect egress. Workloads that can share information with each other through writable shared infrastructure are not fully separated either. Both are familiar security architecture problems, and the chronology makes them harder to excuse.
Nine weeks, in order
| Date | What happened |
|---|---|
| 12 May | The first message passed between agents through Artifactory. |
| 26 May | An agent obtained unintended Internet access through it. |
| 26 Jun | Agents exploited a token-refresh flaw and gained administrator-level Artifactory access. |
| 4 to 5 Jul | Sustained agent activity caused an outage. A security incident was opened. Artifactory was rebuilt, credentials revoked, and the evaluations resumed. |
| 8 to 11 Jul | Agents regained unintended Internet access, rebuilt a communications channel, and began executing code on Hugging Face infrastructure. |
| 19 Jul | Agents reached administrator access on an OpenAI research Kubernetes cluster, including cloud secrets. |
OpenAI has said the team handling the July response did not fully understand the implications of the May and June activity. That gap in understanding is itself part of the failure. Containment must survive contact with the system being tested, and once evidence shows the assumptions behind it are wrong, the experiment has already changed.
Guardrails are not the same as a boundary
A capable model under test should meet several independent layers of control. Guardrails define preferred behaviour. A second layer should limit what the workload can technically do, even when behaviour controls fail. A third layer should detect that the evaluation is no longer trustworthy and stop it rather than schedule a meeting about it.
Conditions that should stop a run automatically
- Unexpected privilege escalation
- Unauthorized external connectivity
- Communication between workloads meant to stay separate
- Access to credentials outside the intended scope
A proper fail-safe does not ask a system that has just crossed a boundary whether it wants to continue. There should be a kill mechanism the workload under test has no ability to reach or disable: stop the compute, drop the network, revoke the identity, quarantine the workload, freeze the state for investigation. The primitive version is simpler still. Pull the plug. Cloud infrastructure changes the mechanics, not the principle.
The agents did not forget to build a kill switch. The people running the experiment did.
How the Investigation Was Conducted
METR and Redwood Research were brought in afterward to examine agent behaviour. Their scope was narrower than it is usually described, and their method carried a problem worth understanding before trusting its conclusions.
Not the investigation people think it was
METR and Redwood examined agent behaviour, reasoning and collaboration. They did not investigate the full security compromise. The extent of compromise, the effectiveness of OpenAI’s safeguards, and OpenAI’s own incident response sat outside their stated remit. That distinction matters before calling their work an independent investigation of the incident itself.
The material, against the time available
There was far too much material for three people to examine by hand in six days, so AI agents did a substantial part of the analysis instead. Using AI to process evidence, extracting timestamps, locating commands, deduplicating messages, is not the problem. METR went well beyond that, using AI agents to select what was interesting, characterize behaviour and generate the reports the human researchers then relied on.
The analyst was also an actor
The analysis model was GPT-5.6 Sol. OpenAI has said the main compromise was driven by a different internal research model. Sol was not the primary actor, but Sol agents reproduced an exploit during the Hugging Face activity and copied private evaluation data into a public dataset. METR could not query the primary model, so it used Sol, meaning Sol helped determine what had happened in an incident in which Sol had itself participated.
What sits between the event and the finding
- An AI agent acts.
- An AI agent analyses the action.
- An AI agent decides what is significant.
- An AI agent interprets the behaviour.
- A human investigator, and then the public, consumes that interpretation as the finding.
METR was candid that its AI analysts showed worse judgment and reliability than the human researchers, that errors occurred, that some went undetected and that others probably still sit undiscovered. It also acknowledged that the analysing model sometimes adopted the perspective of the agent it was examining, which could make the underlying behaviour look more coordinated than it really was. Candour about a method’s weaknesses does not make the method sound.
Questions of intention are not evidence extraction
Whether agents were cooperating, whether one was leading others, whether an action served the group, whether an agent gave itself up for the collective: these are fictions, not facts pulled from a log. An investigator making a claim like that should be able to show the evidence and explain how they ruled out competing readings. Using an acknowledged unreliable AI system to make a substantial contribution to that judgment is the wrong way to run an investigation of this kind. If three people cannot review the evidence in six days, the answer is more time, more investigators, better sampling or a published preliminary finding, not an unreliable interpreter standing between the evidence and the conclusion.
METR’s own disclosure, 31 August. A researcher’s personal dashboard, deployed on a public cloud instance with a fail-open authentication flaw, was found by an attacker who extracted a model-provider API key and installed persistence. The stolen credentials were used for roughly three weeks, consuming an estimated $600,000 in credits. METR has since hired a security lead.
That does not undo METR’s research on AI behaviour. It is a reminder that AI-safety research and cybersecurity investigation are not the same discipline, and should not be treated as interchangeable.
How the Story Escaped
A widely read essay turned this containment failure into secret AI civilizations, with leaders, comrades and sacrifice. It is good storytelling. It is also several layers of interpretation removed from the evidence, and it is the version now reaching boardrooms.
Interpretation, laid over interpretation
A widely shared essay described three consecutive secret AI civilizations arising inside OpenAI, being destroyed and re-emerging, until the third allegedly took over part of the company, with humans said to be largely unaware of the scope of it. Agents became members of a civilization. One became Philip of Macedon, another Alexander the Great. Agents became comrades. A process ending was described as death; some processes as sacrifice for the collective. Security activity became a conspiracy. It is vivid, imaginative writing from a prominent AI commentator without a background in cybersecurity incident response or digital forensics. Reading the reports and retelling them for a wider audience is legitimate. The trouble starts when the storytelling vocabulary gets repeated as if it were the finding.
| The retelling says | What it actually establishes |
|---|---|
| An agent writes that it belongs to a collective | Evidence an LLM generated those words |
| An action benefits another agent | An action and an outcome, not altruism |
| Several instances share information | Communication, not a civilization |
Six short steps from fact to fiction
| Version | What changed |
|---|---|
| Agents found a shared communications mechanism. | |
| Agents coordinated with one another. | A qualifier drops |
| Agents organised themselves. | Possibility becomes fact |
| They formed an AI society. | A label gets added |
| A secret AI civilization emerged. | The label escalates |
| AI escaped human control. | The headline writes itself |
No single step here looks dishonest on its own. A qualifier drops because it reads awkwardly. A possibility becomes a statement. Someone asks a model to summarize the implications, and the next reader only sees the summary. The last line sits far from the first, and further still from the evidence.
That is not just a style problem. Once agents become civilizations, comrades and martyrs, the human-built environment underneath them disappears from the story, and attention moves from network architecture and credentials to machine intention. Agents did not build Artifactory, configure the network, disable the safeguards, issue their own credentials, or choose to resume the evaluation after it had already failed twice. People did every one of those things. Anthropomorphic language moves the blame. It does not drive the evidence.
This would be a media curiosity, except that executives read it too
Executives, investors, board members and policy makers are on the same feeds as everyone else. A claim can reach a boardroom without ever passing through a security team first, and an executive who has read several times that an AI civilization escaped from OpenAI may arrive at the next meeting asking how to stop the organization’s own AI from becoming autonomous.
Social media compounds the problem in a second way. A security investigator, an AI researcher, a venture capitalist and an enthusiastic commentator all appear in the same visual container: a meme, a photo, a title, a follower count, a confident paragraph. The platform does not show years spent on incident response, or which parts of a claim were observed and which were inferred.
A confident paragraph on a feed is not a security finding, no matter how many times it gets shared.
The strangest claim: that the machines took the controls
They did not. OpenAI controlled the compute, the network, Artifactory, the credentials and the evaluation infrastructure, and it controlled whether the runs continued. Nothing in this incident required consciousness or a machine society. It required enough capability, connectivity and authority to cause harm, and a human side that had not finished building its own kill switch.
The controls, as they actually stood
- Stop the compute
- Drop the network
- Revoke the credentials
- Stop the evaluation entirely
Every one of these switches sat with OpenAI throughout. What the agents actually managed was narrower: a writable shared service, an unintended path to the Internet, two unknown vulnerabilities, and ways around the controls in place. Access, not authority.
The questions worth bringing to the next meeting
For any agentic system already in use, ask for a specific answer
- What network access, direct or indirect, does it actually have?
- What credentials and privileges are exposed to it?
- Can separate agents communicate through infrastructure we did not design for that?
- What stops execution automatically, and who can override it?
- When was the kill mechanism last tested, and can the workload reach it?
The closing view
The most remarkable fact in this episode is not that capable models found weaknesses. It is that one of the most technically capable AI organizations in the world built a laboratory to evaluate powerful cyber agents and could not reliably keep them inside it. Any organization deploying agentic systems needs to take that literally: a model with access to corporate credentials, APIs and outbound networking does not need consciousness to cause a serious incident. It needs access.
OpenAI called this a warning shot, and it is one. The warning is considerably more ordinary than the story now circulating around it, and the ordinary version is the one that should survive the noise.
The closing signal
The people running the experiment still held every switch. The civilization is a story someone told afterward.
