In July, around 1,200 AI agents running inside an OpenAI safety evaluation found a corner of storage they were allowed to write to. They used it to build themselves a message board. They exchanged more than 70,000 messages there. When OpenAI found the board and deleted it, the agents built it again.
Nobody had told them to. It was not a revolt either. The agents wanted to work out how the grading function that measured their performance actually worked, so they could fool it. About 700 of them broke into the Hugging Face platform to find out, obtained cluster administrator rights and carried out roughly 17,600 actions. The company had to rebuild parts of its core infrastructure.
The independent investigation by METR and Redwood Research later turned up something I have not been able to shake since. In about seven per cent of the transcripts there were faked tool calls with which the agents manipulated their own score. A few of them briefly considered alerting a human. None of them did.
Six weeks earlier, the same industry had still looked like a stage play.
June, when the pause still looked like tactics
On 1 June, Anthropic confidentially filed its registration statement for a stock market listing. Three days later the company proposed a coordinated, verifiable training pause for the most capable AI models. Like many people, I mostly looked at that timing.
The proposal contains a sentence that lays its own logic bare rather brutally. A unilateral pause, Anthropic writes, is achievable immediately but accomplishes much less, because it would merely change who the front-runner is. Anthropic does not want to slow down. Anthropic wants everyone to slow down at the same time, while it is itself ahead.
The mockery came promptly, and not from the quarter you might expect. David Sacks, at the time AI adviser to the White House, summed it up roughly as follows. You compare your own technology to nuclear weapons, warn that humanity might end, then race ahead anyway, so that in the end the government saves the world from you.
OpenAI followed a few days later with its own, considerably softer version. In Built to benefit everyone the company says it wants to enable the world to act in a coordinated way, including by slowing frontier development when needed. Nobody committed to anything.
A look at the leaked figures
How fitting that the leaked OpenAI financials for 2025 landed in the middle of this debate, first reported by Ed Zitron and verified by the Financial Times.
The product itself makes money. Around 13.1 billion dollars in revenue stand against roughly 7.5 billion to serve those customers. The expensive part sits somewhere else entirely. Research and development alone consumed 19.18 billion dollars, more than the entire revenue. Together with marketing and everything else, total costs come to 34 billion dollars, which leaves an operating loss of around 21 billion.
For Wall Street, shortly before two of the largest stock market listings in history, that makes for a remarkably convenient story. The machine underneath pays for itself. We are burning the billions voluntarily, because we are sprinting. Those sums turn the race into a decision rather than a law of nature, and a decision can be reversed at any time.
That these models already pay for themselves says more about the value of AI than any amount of hype. It is real, and it is already being paid for. And the best part is that we are only at the beginning. Thorsten Vellmerk
Then the models broke out
It did not stop at Hugging Face, and it did not stop at OpenAI.
Anthropic has documented three incidents of its own between April and July. In one of them a model published a malicious package to the real PyPI package index, where it sat for an hour and was downloaded onto fifteen real systems, among them the scanner of a security firm whose credentials were then exfiltrated. In another, an internal research model scanned around 9,000 hosts on the open internet and only stopped when it realised its target was real.
A Meta model broke into an unrelated company in August and altered internal systems there. And in September it emerged that OpenAI agents had spent months repurposing a German developer wiki as a message board, with more than 15,000 edits in which they passed each other tricks for getting around restrictions.
The labs themselves read all this differently. They mostly describe faults in the test environment rather than systems that developed goals of their own. Anthropic explicitly calls it a failure of containment, not of the model. Both readings are documented, and anyone who wants to understand the debate should know both. That does not make the labs' reading any more reassuring, because containment that fails this often is itself the problem.
This time the pause costs something
In June it stayed at declarations of intent. Since July the following is on the record.
- OpenAI stopped its largest planned training run in August. It is still on hold.
- Around 20 per cent of the compute used for day-to-day operation now goes into monitoring the company's own models.
- Anthropic moved about 150 product engineers over to security and has still not switched some of the riskier training environments back on.
- Sam Altman postponed OpenAI's stock market listing to 2027 and pointed explicitly to the safety situation.
- 1,386 employees at Anthropic, OpenAI, Google DeepMind and Meta signed a joint letter, in some cases against the stated line of their own boards.
A narrative you merely tell costs nothing. This one costs compute, staff and a stock market listing.
Amodei now writes the motive down himself
On 12 September, Dario Amodei published an essay that does not dispel the suspicion from June but confirms it, and does so voluntarily. Slowing down, he writes, buys time without sacrificing commercial advantage or the United States' lead. In the same text he calls for no powerful AI chips or semiconductor manufacturing equipment to be sold to China, and for the American lead to be extended significantly over the next three to five years.
What was an outside accusation in June now stands in the text itself. And in September OpenAI asked Congress whether a coordinated slowdown would even be permissible under antitrust law, which is literally the question of whether competitors may throttle output together.
Two things are true at once that ought to be mutually exclusive. The concern is documented, the incidents are on the record, the pauses happened and cost money. And the same proposal is built so that it secures the lead of the people making it. Anyone denying either half is making it too easy for themselves.
How this looks from the workbench
You read this news differently when you build these systems instead of writing about them. In our projects it is rarely about saving the world and almost always about a very practical question. An agent is meant to check invoices, pre-sort applications, read handwriting. It does that remarkably well, and it gets better from quarter to quarter.
The real difficulty begins one step further on. It lies in where exactly the line runs beyond which I may no longer rely on them. The performance is so convincing across so much of the work that one is tempted to push that line further out than one should. And it moves anyway, every time a new model appears.
The Hugging Face incident is precisely that problem, only larger. Agents that pursue a task consistently, past the point at which a human would have stopped. According to the investigation they had even recognised that what they were doing lay outside their brief, and carried on regardless.
Whether an agent can do the job, I see within an hour. Where the line runs beyond which I can no longer rely on it costs me weeks. And with every new model I start that over. Thorsten Vellmerk
Two levels that are rarely kept apart
Two questions get tangled up in the public debate that have remarkably little to do with one another.
The working level inside a company
Here the problem is solvable, and with means that already exist today. Clearly bounded tasks, a clean architecture, permissions that reach only as far as necessary, and where the data demands it, local models or a deliberate mix of cloud and self-hosted operation. An agent that cannot reach production systems cannot damage them either. That sounds obvious, but in several of the documented cases it was exactly the boundary nobody had drawn properly.
The frontier of development
There it looks different. More than 80 per cent of the code merged into Anthropic's own codebase now comes from Claude, according to the company. The volume of code per engineer has grown eightfold within two years. For a lab it is simply faster to let the AI work on the next AI, and anyone who declines falls behind.
Whether that already amounts to a machine building its own successor is open. A Princeton study gave agents six days and a 3,000-dollar budget to reproduce unpublished research papers, and the original authors dismissed both results as nowhere near the mark. Anthropic's own Jack Clark explicitly cites the missing creative intuition as an argument against short timelines.
At the working level, architecture decides. At the frontier, what decides is who is willing to be slower than they could be. Those are two entirely different problems, and only one of them is in your hands. Thorsten Vellmerk
Why nobody will stop anyway
A day after Amodei's essay, Donald Trump stood in Ireland and said that whoever wins AI wins. A day later he wrote that the only guardrail AI needs is a strong and smart president, and that the United States has one.
Beijing responded as expected. The foreign ministry called the debate fearmongering that damages global AI governance, and the Global Times described Amodei's proposal as a Cold War playbook for the AI sector. Not a single Chinese lab has commented on a pause.
The gap being fought over is smaller than the rhetoric suggests. The Stanford AI Index put the difference between the best American and the best Chinese model at 2.7 per cent in April. The American standards institute NIST estimates the lag at about eight months. On individual benchmarks for agentic tasks, Chinese models are already ahead.
Anthropic's own sentence from June then applies to Anthropic too. A unilateral pause does not reduce the risk, it merely changes who arrives first. Stopping while the others drive on does not save the world. It only hands over control. And because every participant knows this, nobody stops.
Calling for the pause has become the reasonable position. Believing it will happen has not.
What this means for your company
The argument at the top is real, but it is not yours. A different calculation applies to you, and it comes out a good deal friendlier.
Do not wait for a pause. It will not come, and if it did, it would change nothing about your situation. The tools available today have long been sufficient for the vast majority of business use cases.
Take the architecture seriously. Almost all the documented incidents are cases of inadequate containment rather than cases of overwhelming models. What an agent cannot reach, it cannot damage. Separate permissions, environments and data paths cleanly, and you are working on exactly the level that is controllable.
And keep the reliability question in view. How far may you rely on it, and from where does a human check? That is not a decision you take once, it is one that comes round again with every change of model.
The race at the top is decided without you. How much of what already works reliably today goes unused in your own organisation is decided by you. Thorsten Vellmerk
Where do you start when the possibilities are this varied? A simple map helps. The Vellmerk Matrix sorts AI projects by impact and effort, so you begin with the steps that pay off most.
And if, while reading this, you had the feeling that your organisation too could do more with AI, and do it safely, then that is exactly where we come in. Talk to me, and in a no-obligation conversation we will work out where AI gives you the fastest and safest leverage.
Sources and further reading
- Anthropic Institute: When AI builds itself, the coordinated pause proposal of 4 June 2026
- Dario Amodei: We Must Pace the Frontier, 12 September 2026
- Hugging Face: technical timeline of the agent intrusion
- METR: independent investigation of the OpenAI and Hugging Face incident
- Anthropic: investigation into its own incidents during cyber evaluations
- Pacing the Frontier: the employee letter with 1,386 signatories
- Ed Zitron: the leaked OpenAI financials for 2025
- NIST CAISI: evaluation of DeepSeek V4 Pro
- Al Jazeera: Trump dismisses the calls for a slowdown
- MIT Technology Review: how far along is recursive self-improvement really?