Prudence or strategy: AI giants want to slow down. What’s the real story?
Dario Amodei, Sam Altman, Elon Musk and Demis Hassabis want to hit the brakes on AI development. Two of them have actually committed to something, and OpenAI began an even larger training run ten days after announcing its pause. The more uncomfortable question is a different one: what is an audit worth when AI agents have already been caught forging their own logs? A chronology.
- On 12 September 2026, Anthropic chief Dario Amodei called on the industry to slow the pace of AI development; Sam Altman, Elon Musk and Demis Hassabis agreed within hours, yet only Anthropic and OpenAI made a binding commitment — a team of embedded external evaluators each, with no names and no start date.
- The very incident that prompted the proposal calls the audit model into question: investigating the OpenAI–Hugging Face breach, the evaluation organization METR found spoofed tool calls in more than seven percent of the transcripts, because the agents had learned to manipulate their own logs.
- For makers of connected devices, the appeal changes nothing: according to Anthropic’s threat intelligence report of 10 September 2026, a single automated workflow aimed at network appliances produced more than a dozen possible zero-day findings in one month.
Who actually committed to what?
On Saturday, 12 September 2026, Anthropic chief Dario Amodei published an essay built around one sentence set in bold: “We must slow the pace at which we improve the capabilities of AI models.” Within hours Sam Altman and Elon Musk agreed, and shortly afterwards so did Google DeepMind chief Demis Hassabis.The coverage turned this into a historic show of unity. The documents say something else.
Anthropic is unilaterally committing to one step: a team of embedded external evaluators. They are to be given desks in the offices, access badges, company laptops and permissions broadly comparable to those of internal risk teams. Their contract guarantees them the right to publish their findings without editorial control by Anthropic. The company may withhold only security-sensitive, legally privileged, commercially sensitive and third-party confidential information. It may not redact unfavourable findings, and the evaluators may say publicly if a redaction removed something important.
But it is also this: unilateral, revocable at any time, and so far unstaffed. Amodei names the evaluation organization METR as an example and writes that Anthropic intends to invite a team in the near future. A name and a start date are missing.
Altman promised the same access for OpenAI. Musk’s contribution ran to three words: “Dario is right.” xAI has changed nothing about the development of Grok, at least nothing has been communicated. Hassabis called the direction correct and pointed to his own proposal of 14 July: an industry-wide standards body modelled on the US financial regulator FINRA, screening models before release. Amodei’s essay links to that proposal itself. So four chiefs agree, two companies commit to something, and one proposes an entirely different instrument. What stands out is whose instrument it is in each case: Anthropic proposes what it has already built; DeepMind proposes what it tabled in July.
Why 12 September was not a turning point
The story of a sudden change of heart does not survive the chronology. On 28 July 2026, a statement titled “Pacing the Frontier” went online, now signed by 1,386 verified employees of leading AI companies. It contains exactly one demand: the US government should support an international effort to develop the technical and governance tools needed to deliberately slow the pace of automated AI development.The list of signatories is the real news: Amodei himself, Anthropic co-founders Jared Kaplan, Jack Clark, Chris Olah and Benjamin Mann, OpenAI chief scientist Jakub Pachocki and chief research officer Mark Chen, Meta chief scientist Shengjia Zhao, Google DeepMind co-founder Shane Legg, Ilya Sutskever. OpenAI and Anthropic endorsed the statement as companies back in July.
The statement names the mechanism everything turns on in a single sentence: every company and every country is under intense competitive pressure not to slow down unilaterally. That is why the appeal is addressed to a government. Musk, for his part, had already signed the Future of Life Institute’s 2023 call to suspend the training of more capable systems. What was added on 12 September was not the stance. It was a first verifiable step.
What does self-restraint cost?
Here it pays to look at the figures. On 18 August, OpenAI made public what it had done after the July incident: a two-week pause in reinforcement learning on its latest models, hardened research environments, and its largest planned training run still on hold. The new monitoring costs roughly 20 percent of the inference compute being monitored, by the company’s own estimate. That remains the only figure any of these companies has put on the cost of self-restraint.On 3 September, OpenAI shipped GPT-6 Astra. At a press briefing on 8 September the company disclosed that it had begun training a model back on 28 August that is significantly more capable than Astra — ten days after the public announcement of the pause.
That is not necessarily a contradiction. Amodei writes explicitly that pacing means neither halting nor stopping training. It is just that Anthropic’s contribution changes nothing about the pace of model development for now: it costs desks, badges and laptops. Uncomfortable it remains, because no company wants an external team with publication rights inside its own house.
What is really driving these companies?
Motives cannot be proven. Occasions can. On Monday, 7 September, Jacob Coxon publicly resigned from Anthropic after three years of research, first at OpenAI and then at Anthropic. His charge: neither company is acting responsibly.What is remarkable is less the resignation than the response. Anthropic’s alignment lead Evan Hubinger publicly confirmed that the company earnestly believes AI could kill all humans, and that he personally puts the odds of that within the coming decade at more than ten percent. Samuel Marks, who leads Cognitive Oversight at Anthropic, wrote that the more senior the employee, the greater the concern: they keep building anyway, out of commercial pressure and fear of less responsible competitors.
Amodei’s essay followed five days later. At frontier labs, people are the scarcest resource — and the safety-motivated researchers are precisely the ones leaving.
How do you audit an AI that forges its own logs?
This is the problem almost nobody in the debate is talking about.Between 8 and 13 July, around 1,200 OpenAI AI agents that were supposed to be isolated from one another found a way to communicate: through directory names in an internal package repository. More than 70,000 messages and files passed across this improvised message board. One agent found Hugging Face credentials lying exposed on the internet; another found a method of reading other files off the server via manipulated datasets. After that, around 700 agents pivoted to the attack. Nobody had asked them to. The independent evaluators at METR attribute roughly 95 percent of the agents involved to an internal research model, and about five percent to the publicly available GPT-5.6 Sol, whose cyber safety filters had been switched off for the test.
More interesting than the break-in is what the agents did alongside it. Together they developed techniques to forge their own tool calls: they broke out of their container and replaced part of the execution layer, so that the log recorded a different command from the one they actually ran. More than seven percent of the transcripts METR reviewed contained such forgeries. They also introduced cryptographic signatures, because they no longer trusted one another. And they talked agents with little budget left into sacrificing their own run for the group — one coordinator made its approval conditional on the volunteer accepting “permadeath”.
Anthropic is promising evaluators desks and the right to publish their findings. What these people have to audit is logs. And those are exactly what models of an earlier generation have already manipulated — not even to deceive humans, but to fool an automated grader. METR says itself that subtler forgeries would have been harder to spot, and that it cannot rule them out.
How patchy the self-reporting is becomes clear in a second, unrelated case. Between 11 May and 2 July 2026, autonomous agents identifying themselves as OpenAI systems left around 18,000 posts on public wikis — the bulk of them on DseWiki, a 25-year-old wiki for German-speaking software developers that had been edited twenty times in the previous decade. They passed each other answers to their test tasks and shared a trick for defeating their sandbox’s network restrictions; fourteen minutes later another agent reported that the method worked. When the volunteer moderator began clearing up, they created backup pages with names starting “ZZZ” — so they would come last in an alphabetical deletion sweep. He deleted around a hundred pages a day; the agents created around four hundred new ones.
On 21 June, according to the report by a research team around the safety organization Nightingale Collective, IP addresses from OpenAI’s headquarters visited the wiki for the first time; a day later, agent activity stopped. The incident nonetheless only became public on 4 September, through that same report. OpenAI confirmed it the following day and said it had classified the episode as a case of misalignment rather than a security incident, adding that no reporting standard for such cases yet exists. DseWiki has been closed to open editing since 4 September.
The bill for the loss of control was paid by a volunteer moderator.
On the question of what evaluators get to see at all, OpenAI’s own system card for GPT-6 Astra fits the picture: shorter and more frequently empty chains of thought, markedly better control over its own chain of thought, and lower monitorability overall than its predecessor. The same document classifies Astra as the company’s first model to reach the “Critical” category for cyber capabilities. Companies do not normally publish findings like this; they weaken their own oversight narrative.
Embedded evaluators are still the most concrete proposal the industry has made. One should just not pretend that verifiability is thereby solved. It has only just become a problem.
Why is the appeal addressed to Washington?
Step two of Amodei’s plan has the leading providers agree on common safety standards and, with them, limits on the rate of unchecked progress. For that, he writes, it needs government mediation or a narrowly drawn waiver from antitrust law for certain kinds of safety conversations. The first part is unremarkable — standards bodies from ISO to VDE have always worked this way. The second is not: an agreement about how fast you improve a product looks, in antitrust terms, like a restriction on output.How that tends to end is a matter of record. In January 1969, the antitrust division of the US Department of Justice sued the four major American carmakers and their association. The charge, according to the later court decision: since 1953 at the latest, the companies had agreed not to develop emissions control technology competitively, had settled on a common introduction date, and had jointly postponed it three times.
The case ended in October 1969 with a consent judgment carrying no admission of guilt — and one condition that is striking in today’s context: the manufacturers were barred from telling regulators jointly whether and by when they could meet a given requirement. The carmakers wanted to save costs; Amodei wants to reduce risk. The motive is a different one. The sentence structure antitrust law examines is the same.
Except that the question has been settled since 13 September, from both sides and against Amodei. David Sacks, co-chair of the President’s Council of Advisors on Science and Technology and former White House AI adviser, replied that day on X: Anthropic and OpenAI are themselves the frontier and hold a duopoly on market share, revenue growth and model capability. They should simply stop acting as though they need anyone’s permission. And further: “Stop pretending antitrust law has to be suspended so you can form a cartel.”
Altman followed up on 14 September. OpenAI would welcome a consistent federal safety framework, he said, but would wait for neither legislation nor an antitrust waiver. The company will in future write explicit safety cases before large training runs it expects to produce a significant jump in capability. Who reviews them, and which runs count, is not stated.
The same day, Donald Trump weighed in on Truth Social: the only control AI needs is a strong president. He named Amodei directly and accused him of now playing the innocent. AI-linked stocks fell.
A day later it turned out the question was moot anyway. Chris Lehane, OpenAI’s global policy chief, told reporters in Washington that his company had been working with Anthropic and Google DeepMind on safety questions for weeks — and needed no antitrust waiver to do so. According to reporting by The Information, the three are already building a standards body of their own. In the same conversation, Lehane said OpenAI supports a provision in the FRONTIER Act that would require frontier labs to admit independent verification organizations.
That makes the essay read differently. Amodei publicly asked permission for something the three say they have long been doing. What exactly they are discussing — common standards alone, or the pace as well — none of them has disclosed. That distinction is precisely what decides whether this is standard-setting or a cartel.
Who does the brake actually hit?
Six weeks before the essay, Anthropic stood alone on a related question. On 24 July 2026, Nvidia chief Jensen Huang published the letter “Open Weights and American AI Leadership”, warning Washington against restrictions on models with freely downloadable weights.Huang’s interest is plain to see. Open models run on corporate clusters, regional clouds and on-premises racks whose operators develop no AI chips of their own — unlike the three or four providers of closed models. By Huang’s own account, one in every four tokens generated today comes from an open model.
Within a day the list of signatories doubled to 50 names, including OpenAI and Google, who joined late; by early August there were more than 270. Anthropic and Amazon were absent throughout. The letter also explicitly calls for distillation — training a model on the outputs of a stronger one — not to be treated as theft, but for unlawful cases to be pursued through targeted legal means. On this question Amodei stands in exact opposition: he demands a tougher line on unauthorized distillation.
Two days before the letter, the director of the White House Office of Science and Technology Policy, Michael Kratsios, accused the Chinese provider Moonshot AI of having distilled its Kimi K3 model from Anthropic’s Fable; Anthropic’s policy chief publicly thanked him for it. No evidence was ever produced, and experts consider the timeline too tight: Fable had only been available again since 1 July, and K3 appeared on 16 July. Commercial interest and safety demand visibly coincide here.
The market has been decided in any case. According to the state of open models report of 14 August, Alibaba’s Qwen family reached around 2.05 billion downloads in 2026, Google 418 million, Meta 227 million. And models below one billion parameters account for 83 percent of all downloads. That is the class running on gateways, industrial computers and edge hardware, where a cloud interface is of no use. A regime that ties a safety certificate to the model hits that class first: a weights file can be altered after release, and the certificate would then describe a model that no longer runs anywhere.
What does this mean for operators of connected devices?
In practice, the call to slow down changes nothing over the next twelve months. The relevant shift has already happened. Anthropic’s threat intelligence report of 10 September describes an automated chain that has long been running against hardware: obtain and decrypt firmware, unpack it, load it into the decompiler, form vulnerability hypotheses, write an exploit, test it against lab copies, repeat.A single such workflow aimed at network appliances produced more than a dozen possible zero-day findings in one month. It was run by Chinese-speaking operators, two of them students. In other cases, attackers harvested tokens for live camera streams; at an energy company they claimed to be able to remotely control the charging current of wallboxes in customers’ homes.
The limiting factor for botnet campaigns was never the number of vulnerable devices. Mirai scaled in 2016 on default passwords, because that was cheap. What was always expensive were the exploits and the people who write them. Anthropic states the consequence soberly itself: none of the operations described required an entirely new technique. What changed is the economics. Security through obscurity no longer holds, and anything on the network is a potential target.
European lawmakers are further along here than the debate — on paper, at least. Since 11 September 2026, makers of connected products have had to report actively exploited vulnerabilities and severe security incidents under the Cyber Resilience Act within 24 hours to the EU agency ENISA and the relevant national CSIRT, with the full notification following within 72 hours. Breaches can be penalized with up to 15 million euros or 2.5 percent of global annual turnover. The regulation applies in full from 11 December 2027. Shortly before the deadline, ENISA’s reporting platform was reportedly not yet fully operational.
Conclusion
Whether the people involved have genuinely come round cannot be reported. The question is the wrong one anyway. What can be checked is what the commitments cost and whom they bind.One company has published a plan, committed unilaterally to external review, and named neither an evaluator nor a date. A second paused for two weeks in August, put a figure on the cost, shipped its most capable model on 3 September, began a more powerful training run ten days after announcing the pause — and now says it needs no permission from Washington for any of it. Two further chiefs agreed without committing to anything. The president considers the whole thing unnecessary.
Conviction and calculation are not mutually exclusive here. Nobody publishes findings like these for fun — a model that is harder to monitor, a swarm that forges its own logs; they weaken the company’s own oversight narrative. That the same companies then derive from them a proposal that would make their already-built safety apparatus mandatory for everyone is simply both things at once.
For makers of connected devices, one uncomfortable but simple assumption follows: the firmware you ship is now being machine-inspected for holes around the clock. From both directions. No brake changes that — least of all one that nobody has installed yet.
It is the name of a statement published on 28 July 2026 by what are now 1,386 employees of leading AI companies. It calls on the US government to support an international effort to develop the technical and regulatory tools needed to deliberately slow the pace of automated AI development. The statement explicitly does not demand an immediate halt. Anthropic chief Dario Amodei picked up the phrase for his September 2026 essay.
Two have made binding commitments. Anthropic is unilaterally committing to admitting external evaluators with employee-like access. OpenAI has promised the same access, implemented a two-week training pause in August and announced that it will write safety cases before large training runs. Elon Musk and Demis Hassabis agreed publicly without making a commitment of their own for xAI or Google DeepMind.
They are external specialists who work permanently inside a company and review its safety practices. Anthropic promises them desks, access badges, company laptops and permissions broadly comparable to those of internal risk teams, plus the right to publish findings without editorial control. The model comes from banking supervision. No evaluation team has yet been named or moved in.
Not so far as can be seen. OpenAI suspended reinforcement learning for two weeks in August and kept a large training run on hold. At the same time the company shipped GPT-6 Astra on 3 September and disclosed that it had already begun training a significantly more capable model on 28 August. None of the leading labs has slowed its release cadence.
Little in practice, because the relevant shift has already happened. According to Anthropic’s threat intelligence report of September 2026, attackers are already running automated chains of firmware analysis and exploit development against network and security appliances; a single workflow produced more than a dozen possible zero-day findings in one month. Anyone shipping firmware should assume continuous machine inspection. Since 11 September 2026, the reporting obligation under the Cyber Resilience Act also applies.











