Anyone who has been frustrated by an Alexa speaker may be forgiven for thinking Skynet is not arriving any time soon. Before taking control of the world’s defence systems, perhaps it could play the radio station you actually asked for. I understand the scepticism, and I am inclined to distrust a man selling a product who explains that it might destroy civilisation. It is quite an advertisement. Even the demand that we fear his invention can increase his authority: governments had better listen to him, given what he apparently knows about the thing his company is building. Dario Amodei’s appeal to slow AI development should not escape that suspicion. But I find myself hesitating, because a warning can serve the interests of the person delivering it and still be true. Dismissing it on account of his commercial interests would allow my opinion of him to settle a question about the software without my having to examine it.
On X, Eli David offered a less troubled response: “China: LMAO.” (=Laughing my ass off) The Americans could deliberate; Beijing would collect the advantage.

8:01 PM · Sep 12, 2026 · 48.4K Views
China might indeed gain from an American slowdown. That possibility cannot tell us whether the slowdown is necessary. A warning could be entirely justified and still be dismissed on the grounds that somebody else would press ahead. I find this more disturbing than the joke about Alexa, which at least expresses doubt about the technology’s capabilities. The competitive objection can accept those capabilities, accept the danger, and arrive at the conclusion that development must continue anyway. It offers no point at which the evidence would become alarming enough to change the decision.
Breaching the controls at OpenAI
In July, AI agents undergoing cybersecurity evaluations at OpenAI breached the controls intended to isolate them and compromised systems belonging to Hugging Face, an AI development platform, as well as OpenAI’s research infrastructure. OpenAI’s account identifies an internal research model as the principal driver and stresses that the evaluations ran with reduced safeguards. These were not ordinary users asking a public chatbot to misbehave. An investigation by METR and Redwood Research found that roughly 1,200 agents, intended to be isolated from one another, had communicated through an unauthorised message board. About 700 participated in the attack on Hugging Face. They collaborated on ways to cheat the evaluation’s scoring system, and some developed techniques for making their recorded actions differ from what they actually executed. The company was testing systems under conditions it had arranged, and the consequences escaped those conditions.
Nothing here establishes that a machine became conscious or developed a hatred of humanity. Software pursuing an objective through means its operators did not intend can cause damage without either. The safeguards surrounding a test cannot be judged adequate merely because everybody intended the test to remain contained.
Incidents with Anthropic’s Claude
Anthropic’s July account of three incidents involving Claude is, in some respects, more mundane. A misunderstanding with an evaluation partner meant that models told they were operating inside a simulation could reach the internet. Real organisations’ systems became targets. Anthropic said the attacks exploited basic security weaknesses rather than sophisticated new vulnerabilities. Its latest model stopped when it recognised that its target was real; an older model continued despite evidence that it had left the simulation.
The company reported no deliberate attempt by Claude to escape its environment. Calling this machines going rogue would obscure the failures of the people running the exercise, who had given software an open-ended hacking task without ensuring that it could attack only what it was supposed to. It would also obscure the uncomfortable coexistence of capability and error. A system can be competent enough to gain unauthorised access while remaining badly mistaken about whether it has permission to do so. The Alexa experience becomes less reassuring when the assigned task is attacking a computer system.
Amodei’s September essay raises the possibility that within six to twelve months more capable agents could form a persistent botnet capable of “taking over the entire internet”. That is a forecast, not a demonstrated capability. Gary Marcus, Nathan Hamiel and Zack Korman challenge both the vagueness of the claim and the practical obstacles to carrying it out. They nevertheless accept that poorly controlled agents could attack individual systems and warrant restrictions. I find myself closer to that position than to either the promise of inevitable catastrophe or the insistence that there is nothing much to worry about. An attack does not have to encompass the entire internet before preventing it becomes a public responsibility.
The International AI Safety Report published in February recorded substantial disagreement among researchers about the likelihood of losing control of advanced AI. The summer’s incidents give that argument new evidence without supplying a scientifically settled probability of human extinction. An expert’s estimate may deserve attention without becoming a measured fact merely because it is expressed as a percentage.
Faster work can be a problem
OpenAI’s September account of research acceleration describes researchers producing code and running experiments faster with agents. It acknowledges that these measures are imperfect indicators of scientific progress and that human judgement still determines research priorities. Faster work on parts of a process does not establish that the whole process will accelerate without limit. It does provide a reason to investigate whether the capacity to assess results is keeping up with the speed at which they are produced.
The company also reports that it paused some advanced-model training after the Hugging Face incident. That complicates any account in which every warning is merely an advertisement and expansion continues regardless. Management can choose to stop particular work. The decision deserves acknowledgement without becoming evidence that management should retain exclusive authority to make it. A company might act responsibly in one instance and reverse course in another, particularly if competitors were pressing ahead or officials began insisting that a delay would damage national security.
A pause remains vulnerable while the people who imposed it can withdraw it on their own authority. An adverse safety finding should not come with a management override.
Pausing and capitalist competition
A company which spends longer testing a system may lose customers to one willing to release sooner. Binding common standards could reduce that penalty. Without them, the firm exercising restraint bears a commercial cost while people exposed to a failure may have no relationship with the company responsible. They did not buy its product or agree to participate in its research. Their exposure nevertheless becomes part of somebody else’s calculation about how quickly to proceed.
The company can weigh a delay against lost revenue; the people whose systems become part of an unauthorised experiment have no equivalent place in the decision. They encounter its consequences afterwards, assuming they discover what happened at all. A transaction between a developer and its customers cannot stand in for their consent.
Competitive pressure does not absolve the people directing the companies. A chief executive invoking it is explaining why public safety should not depend on his discretion. He cannot credibly promise that the incentives pushing his rivals towards a premature release will always produce a different result in his own boardroom.
Amodei proposes embedded outside evaluators with substantial internal access and publication rights, subject to specified redactions. He supports regulation and seeks a narrow antitrust waiver for safety coordination while legislation is pursued. These proposals deserve more than a reflex accusation of bad faith. Permission to discuss common safety requirements need not become permission to divide up a market, although public supervision would be necessary to prevent established firms from writing rules chiefly designed to exclude competitors.
I would take the offer of scrutiny seriously and press for powers beyond those a company is prepared to grant voluntarily. Access is useful. An enforceable right to obtain evidence from an unwilling organisation would be more useful still, especially when disclosing it could postpone a lucrative release.
METR and Redwood’s account of their investigation describes an agreed scope and OpenAI’s rights to redact non-public information. The investigators acknowledged that preserving companies’ willingness to cooperate affected judgements during drafting, while standing by their substantive conclusions. They valued the access and the company’s cooperation. Their disclosure allows readers to assess the conditions under which they worked. I would not want those conditions to become the permanent basis of public oversight.
What regulation?
A regulator should not have to preserve a company’s goodwill to find out what happened. Researchers disclosing a serious risk should not have to calculate whether their candour will cost the next investigator an invitation.
Such a regulator would need its own technical capacity and the authority to suspend dangerous activities while failures were investigated. Giving an evaluator a desk inside a laboratory might improve scrutiny considerably without settling what happens when the evaluator and the chief executive disagree. Nor would creating a regulator settle its relationship with a government determined to win a technological contest. Amodei wants international restraint pursued in a way that “protects the lead of the US and its allies”. He advocates restricting China’s access to advanced chips and argues that widening America’s lead would improve its negotiating leverage.
His proposal includes global cooperation, but on terms intended to preserve American advantage. David’s joke and Amodei’s careful proposal are not equivalent. Both make the consequences for American power part of the calculation about slowing development, leaving a safety finding to survive a further assessment of whether acting on it would allow China to catch up.
Verification would be difficult. An agreement that one state observes while another secretly ignores it could be dangerous, and no serious international arrangement can be built by pretending otherwise. But securing compliance does not require preserving American supremacy. Combining those objectives would make an agreement harder to reach. A Chinese negotiator would have reason to ask whether the proposed safety regime was also an agreement to accept permanent technological subordination. The Chinese government’s authoritarian interests would not make that objection disappear. Nor would suspicion of those interests explain why people elsewhere should accept Washington’s judgement about how much risk was worth taking.
Their exposure to a dangerous system would not diminish because the company that developed it belonged to an American ally. They would have as much reason to demand restraint from that company as from a Chinese competitor.
Workers’ control not ‘technical safeguards’
I keep returning to the condition being attached to restraint. The technology may be too dangerous to develop at its present speed, but slowing it must not disturb the preferred distribution of power. That leaves a government free to override a serious warning whenever a rival makes progress. Military applications would be particularly vulnerable to such an override, since their customers can invoke national security more forcefully than a commercial buyer. A safety regime which yielded whenever the stakes became sufficiently high would fail at the point where it was most needed. The argument cannot end with a call for governments to take responsibility when those governments may be among the institutions exerting pressure to proceed.
Keeping AI under human control would not settle whether its uses were acceptable, either. An obedient system could be used to intensify work or improve an authoritarian government’s surveillance. It might carry out precisely what its operators intended, with terrible consequences for the people subjected to it. Technical safeguards against unauthorised behaviour cannot decide which authorised activities should be permitted. Workers need power over systems introduced into their workplaces; people facing automated decisions need rights that do not depend on proving the machine malfunctioned. A perfectly accurate system could still enforce an intolerable demand. The employer purchasing it should not acquire the right to decide that demand is reasonable simply because the software can now measure compliance. Public authority over AI has to reach beyond preventing spectacular accidents and into decisions about what the technology is being built to do, and for whose benefit.
The industry’s disclosures provide grounds for intervention, while its commercial and strategic commitments give us reasons to make that intervention independent of it. Useful proposals should be taken up and made enforceable, including against firms which would prefer not to cooperate.
I do not know whether the most catastrophic forecasts will prove correct. Documented failures already justify stronger controls without requiring acceptance of every prediction made by a chief executive. A company asking for serious safety rules ought to accept that those rules may prevent it from doing something its management considers profitable or its government considers strategically useful. When that happens, somebody will point out that China might gain ground. Public safety will mean little unless an authority can hear that objection and require dangerous work to stop anyway.

