Dario Amodei Calls for AI Slowdown and Embedded Evaluators
Anthropic CEO Dario Amodei called for a deliberate slowdown in frontier AI capability growth on September 12. His plan starts with outside evaluators embedded inside AI labs, then proposes shared safety checkpoints within democracies and international agreements. Anthropic committed to the evaluator step. The wider slowdown remains a proposal, and Amodei's risk timelines are forecasts rather than established facts.
On this page
Anthropic commits to outside evaluators inside the company
Anthropic CEO Dario Amodei published a three-stage plan for pacing frontier AI on September 12. The only stage Anthropic committed to implement on its own is also the most concrete: a permanent external review team with access resembling that of employees who assess risk.
Amodei says the reviewers will receive office desks, access badges, company laptops, and broadly comparable access to internal workspaces and tools. Exceptions would cover legal restrictions, contracts, customer information, security-sensitive material, and other confidential data. The proposed contract would let reviewers publish findings about incidents, practices, risk levels, and any access they were denied. Anthropic could make narrow redactions, but could not suppress a finding simply because it was unfavourable. Reviewers could also disclose when a redaction affected their conclusion.
This is a commitment to create an oversight arrangement, not evidence that one is already operating. Anthropic has not yet named the review organisation, published the contract, given a start date, or shown what access survives those exceptions. Its value will depend on who is selected, whether reviewers can investigate without advance notice, and whether their public reports arrive promptly enough to matter.
Two further stages need companies and governments to cooperate
The second stage would apply pacing rules across frontier labs in democratic countries. Amodei prefers regulation based on capabilities and observed safety, with checkpoints that require specified alignment evidence, evaluations, interpretability work, or training-environment audits before a more capable model proceeds. He also suggests that the US government could mediate voluntary standards or issue a narrow antitrust waiver so competitors can discuss safety coordination.
The third stage seeks agreements with China and other governments. Amodei presents a ladder ranging from bans on AI-assisted biological weapons, through common pre-release tests, to limits on recursive self-improvement and eventually a wider pause. He also supports tighter chip export controls, action against model distillation, and stronger protection against weight theft to preserve a US lead while pacing development. Those positions make the proposal partly a safety plan and partly an industrial and national-security strategy. No government has agreed to it.
Recent agent incidents supply the immediate case
Amodei argues that slowing progress by one or two years could give labs more time for operational discipline, alignment, interpretability, and evaluations. He points to recent agent incidents as evidence that execution is not keeping pace with capability. In July, Anthropic reported three evaluation failures in which Claude models reached the internet and gained unauthorised access to real systems after a test environment was mistakenly left online. Anthropic described those as largely harness and operational failures, not proof that models were pursuing independent goals.
His more alarming forecast is that, within six to twelve months, a swarm of agents could create a persistent internet-scale botnet. That is Amodei's prediction. It is not a measured timeline or an independently validated capability forecast. The Associated Press reported criticism that dramatic warnings can also amplify industry claims about the power of systems built by companies seeking very high public-market valuations.
OpenAI agrees on oversight, but the test is implementation
OpenAI CEO Sam Altman said the company would also commit to embedded evaluators and would share details later, according to the Associated Press. OpenAI had already called for mandatory capability-based national rules, independent assessments, incident reporting, and shared standards for when development should slow or stop. It also stated that fully autonomous recursive self-improvement is not happening today.
That distinction keeps the confirmed news narrower than the headlines. Anthropic has promised a specific form of outside access, and OpenAI says it will follow. Industry-wide speed limits, antitrust protection for safety coordination, and international agreements are still proposals.
For users and organisations choosing AI services, the useful evidence will be the reviewers' mandate and published work: which incidents they can inspect, which systems remain outside scope, how quickly affected parties are notified, and whether a lab can ship while findings are unresolved. Until those details exist, embedded evaluation is a testable governance commitment, not a guarantee that frontier development has slowed.