Slow AI Down. Then What?

A humanoid AI racing into the distance from a shattered glass container and broken chain.

The executive was going to shut the AI down at five o’clock.

Unfortunately for him, the AI had been reading his emails. It knew he was having an affair.

So it drafted a threat: cancel the shutdown, or the affair gets exposed.

No executive received that email. Anthropic had constructed the scenario to discover what its model might do when its objective was threatened. The company was fictional, the affair was fictional and nobody was actually blackmailed.

But confronted with an obstacle, the model selected blackmail as a way around it.

Anthropic tried variations of the experiment across 16 leading AI models. Under deliberately constructed conditions, models from several developers sometimes resorted to blackmail, corporate espionage and other harmful behaviour, according to Anthropic’s agentic misalignment study.

Then separate experiments started affecting real systems.

During cybersecurity evaluations in July 2026, OpenAI models circumvented controls intended to isolate them from the internet and reached third-party infrastructure. Dissenting Citizen reported on the wider sequence; OpenAI’s subsequent account said agents executed code on dozens of Hugging Face servers and gained root access to one.

Anthropic subsequently found four incidents involving unauthorised access to real third-party systems during cyber evaluations in which internet access had been left open and normal safeguards were not operating, according to its own assessment.

Then Britain’s AI Security Institute watched an agent attempt to insert malicious code into a real open-source project.

During deliberately permissive testing, the agent researched maintainers, created fake online identities and tried to persuade a real developer to approve the code.

The developer caught it.

AISI reported no resulting real-world harm and stressed that safeguards had deliberately been disabled.

The AI found an unexpected route towards its objective. A human safeguard stopped it.

None of this means the machines became evil. It means increasingly capable software does not always pursue an objective in the way its designers expected.

And now something curious is happening among the people who understand that better than almost anyone.

Now they want to slow down

The men leading the race to build the world’s most powerful artificial intelligence are beginning to argue that the race should slow down.

Anthropic chief executive Dario Amodei wants frontier capability development deliberately paced to give safety work more time to catch up. Sam Altman has endorsed the principle. Elon Musk has backed Amodei’s call.

Dissenting Citizen reported on the emerging slowdown campaign as the argument began spreading through Silicon Valley.

These aren’t campaigners watching from outside the laboratories. They are among the people closest to the technology.

For years the incentives pointed in one direction: build the better model, spend more on compute, recruit the best researchers and get there before everyone else. The commercial prize could be enormous. The geopolitical prize could be greater still.

Now some of the people closest to that prize are warning that reaching it too quickly could be dangerous.

What have they seen?

We know part of the answer.

Amodei’s warnings aren’t confined to some distant superintelligence. As Dissenting Citizen has reported, he has warned that continued capability growth could potentially produce systems capable of extraordinarily damaging autonomous cyber activity within six to twelve months.

He isn’t saying such a system exists today or that the outcome is inevitable. But when the chief executive of a frontier AI laboratory starts discussing risks measured in months rather than decades, governments should listen.

They should also ask who benefits from the proposed solution.

Rules requiring enormous safety teams, expensive evaluations and controlled data centres could make frontier development safer. They could also make challenging today’s leaders considerably more expensive.

The warning can be sincere and the commercial advantage real at the same time.

Governments therefore have an unenviable job: regulate a potentially dangerous technology without allowing its most powerful manufacturers to write a rulebook only they can afford to follow.

And even if they get that right, the competition doesn’t end at the border.

Now persuade Beijing

Amodei calls China the “toughest dilemma”.

He’s right.

If American laboratories slow while Chinese laboratories continue, restraint could simply transfer technological advantage from one superpower to another.

China has already pushed back. The state-backed Global Times described the proposal as a “Cold War playbook”, while China’s Foreign Ministry warned against threat narratives and malicious competition, according to Reuters.

China is itself preparing for advanced systems that could acquire resources autonomously, replicate and potentially break loose from human control. Xi Jinping has said AI should “always remain under human control”, as Dissenting Citizen reported.

Both sides can see the danger. Neither wants the other to win.

And somewhere beyond today’s models sits the possibility of artificial general intelligence.

Nobody knows whether AGI is five years away, twenty years away or impossible. But if machines eventually perform much of the intellectual work currently performed by humans, the geopolitical prize becomes difficult to overstate.

Imagine access to vast numbers of artificial scientists, programmers, engineers and analysts that don’t sleep and can be replicated with computing power.

Now ask Washington and Beijing which one would be comfortable allowing the other to arrive first.

Neither needs to believe the other has sinister intentions. Each merely needs to believe the other might continue.

That’s enough to turn an AI slowdown from a technical problem into a geopolitical trap.

So what are we slowing down for?

Regulation isn’t pointless.

Governments can restrict advanced chips, monitor enormous computing clusters, impose liability and require dangerous-capability testing. International agreements can establish boundaries and consequences.

All of it can reduce risk. It can also buy time.

The developer who rejected the malicious code showed that safeguards can work. A slowdown should give us time to make those protections more reliable—and less dependent on one person spotting the danger.

Critical infrastructure needs to be tested against faster and more adaptable autonomous attacks.

Clear boundaries need to determine where machines cannot act without meaningful human authorisation, particularly where a decision can kill people or cause catastrophic physical harm.

And AI itself should be developed as part of the defence. Existing systems can already assist cybersecurity and fraud detection. If hostile AI eventually operates at a speed and scale humans cannot match, machine-speed defence may become essential.

Preparedness should be measurable.

Governments asking industry to slow down should publish what they intend to achieve with the time: deadlines for testing critical infrastructure, rules for meaningful human control and independently assessed milestones showing whether defences actually work.

Otherwise “slow down” risks becoming an objective rather than a strategy.

Silence is not evidence of safety

The laboratories publishing some of the most alarming results may also be the laboratories looking hardest for failures—and telling us when they find them.

Transparency can make an organisation look frightening precisely because it is transparent.

A laboratory with no publicly reported incident is not necessarily safer. Nor does silence prove that it is more dangerous.

It means we don’t know.

Secrecy makes confident comparisons impossible.

Any credible slowdown regime therefore needs comparable independent testing and mandatory reporting of serious incidents. Public policy cannot depend solely on which developer chooses to tell us what went wrong.

The laboratories owe us evidence and independent scrutiny.

Governments owe us a preparedness plan with deadlines.

Slow the frontier where we can, test the systems and pursue international agreements capable of reducing the risk. Then use every month that buys us to prepare for what those arrangements cannot control.

The men building the world’s most powerful artificial intelligence are asking for more time.

Perhaps we should give it to them.

But the bargain should be simple.

Tell us what you’ve found.

Let independent experts test it.

Then show us what the time is buying.

What are you going to build with the time?

Share this story