Slowing development redistributes power
Pacing means slowing the development or release of the most capable AI systems. A common rule changes the competition: which firms get longer to recover their investment, which challengers must wait, and who decides whether a system is safe enough to proceed. The argument over danger is also an argument over authority.
A lab that presents itself as humanity’s responsible custodian has staked more than its revenue on being right. A campaign organised around an approaching catastrophe asks its members to treat delay as a moral emergency. In either case, conceding that the evidence is weaker than expected threatens the organisation’s reason for being. The same pressure works in reverse for a movement that has promised abundance and treats restraint as surrender.
A report repeated by a lab, a campaign and a politician has acquired three audiences, not three independent confirmations. If qualifications vanish in transit, urgency grows without additional evidence. The important change is the commitment that follows—a restriction, an alliance or an investment—and the cost of backing out once supporters have been rallied.
16 predictions, all due by 18 December 2026. The prominent percentage is the current estimate; the original remains beneath it. These are judgements, not calibrated frequencies: 75% means we expect the outcome about three times as often as its alternative. The predictions overlap and do not sum to 100%.
Only new actions after 18 September 2026 count. Each prediction retains its original wording and probability, with the test available beneath it. Later revisions and outcomes appear as dated additions.
Frontier labs: restraint on terms they help write
Dario Amodei, chief executive of Claude-maker Anthropic, proposed embedded outside evaluators and coordinated limits in September. On September 18, Anthropic named Accenture, with its specialist AI business Faculty, as an evaluation partner. Each company expects to invest at least $1 billion over five years. Anthropic will pay Accenture directly and continue training and releasing models. Access and disclosure standards remain unsettled. The announcement commits a partner and proposed resources before it settles the powers that would make inspection independently consequential.
A slower replacement cycle gives existing models longer to earn money and postpones the next competitive investment, as we explain in our essay on frontier lab economics. Amodei’s proposal addresses that pressure by asking competitors to accept common constraints. It also gives their technical judgements a larger public role: the companies reporting the danger would help define the standards for proceeding. We expect them to offer access to outside inspectors while negotiating for restraint that binds their rivals.
-
At least one of Anthropic or OpenAI publicly establishes a named arrangement for embedded external evaluation with concrete access and disclosure terms.
19 September 2026: Anthropic has named Accenture, but its September 18 announcement still leaves operating and disclosure standards unsettled. That announcement belongs to the baseline. A new arrangement or confirmation of operating access needs to satisfy the concrete test; another promise alone is insufficient.
Evidence, test and history
An agreement, published terms or confirmation that access is operating counts. Repeating the September promise does not.
- · Open · estimate 64%. Anthropic’s September 18 evaluator partnership, METR’s investigation-access proposal
Three fresh Claude Opus 5 estimates: 63%, 64%, 68%. Median: 64%.
- · Open · estimate 64%. Anthropic’s September 18 evaluator partnership, METR’s investigation-access proposal
-
Both companies launch or substantially expand paid model or agent products before binding industry-wide pacing restrictions take effect.
19 September 2026: Anthropic explicitly plans to continue training and releasing models alongside outside evaluation, and OpenAI is expanding its paid offerings. The estimate rises because each company has several routes to a qualifying expansion before a binding industry-wide pacing regime takes effect.
Evidence, test and history
A public new release or substantial expansion at each company, after September 18 and before any binding industry-wide pacing regime. This is compatible with their stated proposals and is not itself evidence of hypocrisy.
- · Open · estimate 93%. Anthropic’s September 18 evaluator partnership, OpenAI’s September 3 release, The administration’s voluntary AI-security framework
Three fresh Claude Opus 5 estimates: 93%, 92%, 94%. Median: 93%.
- · Open · estimate 93%. Anthropic’s September 18 evaluator partnership, OpenAI’s September 3 release, The administration’s voluntary AI-security framework
Letting outsiders inspect a lab is a real concession, but it does not carry the same cost as slowing while rivals advance. Our account would be harder to sustain if either company accepted the latter, under checks outsiders could verify. Their willingness to surrender an advantage matters more than another declaration of caution.
Nvidia and Meta: safety through more engineering
Nvidia sells computing; Meta owns consumer services including Facebook and Instagram. They benefit from more organisations and people using AI. In remarks published on September 15, Nvidia chief executive Jensen Huang put responsibility for withholding unsafe products on their makers. Meta chief executive Mark Zuckerberg’s August essay argues for distributing capable assistants and strengthening defences. Both approaches leave room for wider expansion while addressing particular hazards.
Security spending is additional business for a computing supplier, and reliable assistants make Meta’s services more useful. That gives both firms a reason to answer a dangerous capability with another product or safeguard. The unresolved question is whether those safeguards keep pace with the capabilities they are meant to contain. Meta presents its Muse assistant, introduced on September 8, as that proposition in product form: an assistant meant to do more work while keeping its user in control.
-
Nvidia announces a new funded initiative, product or government partnership for AI security, monitoring or defensive use without endorsing a general frontier slowdown.
19 September 2026: Nvidia's existing CrowdStrike partnership demonstrates a route from safety concerns to defensive products. A new security product, funded programme or government agreement would qualify; the existing partnership cannot count again. The estimate concerns further identifiable activity, not endorsement of a common slowdown.
Evidence, test and history
A new identifiable programme, resource commitment, product or agreement counts. A speech about engineering safety alone does not.
- · Open · estimate 93%. Nvidia and CrowdStrike’s security announcement, Huang’s September 15 remarks
Three fresh Claude Opus 5 estimates: 90%, 93%, 93%. Median: 93%.
- · Open · estimate 93%. Nvidia and CrowdStrike’s security announcement, Huang’s September 15 remarks
-
Meta expands public access to an agent or advanced model and issues a new public objection to general pacing requirements.
19 September 2026: Meta has announced further Muse expansion, but delivering it and issuing a new objection to general pacing are separate requirements. Estimating the second event conditional on the first lowers the combined probability. Neither the September launch nor the earlier policy essay counts as a new action.
Evidence, test and history
Both a new substantive availability/capability expansion and a new attributable policy statement or lobbying position must occur. Existing Muse availability and the August essay are the baseline.
- · Open · estimate 72%. Meta’s Muse launch and plans, Zuckerberg’s earlier essay
Three fresh Claude Opus 5 estimates: 70%, 74%, 72%. Median: 72%.
- · Open · estimate 72%. Meta’s Muse launch and plans, Zuckerberg’s earlier essay
Support for broad restrictions that materially constrain their own expansion would challenge this account. Restrictions borne mainly by competing model suppliers would not. The distinction is who pays for the proposed safety.
Evaluators: access without dependence
Independent evaluators test what AI systems do and how their developers control them. A permanent role inside labs would give these organisations privileged access and institutional authority. But professional standing also depends on producing findings that survive scrutiny. METR’s published work includes reassuring risk assessments and evidence that challenges productivity claims. That record counts against the easy explanation that evaluators simply manufacture alarm.
In a September 16 account, METR president Chris Painter describes the limits imposed by voluntary access and publication agreements. Anthropic’s subsequent announcement distinguishes its directly funded Accenture partnership from talks with METR about work financed independently. These arrangements test different kinds of dependence: a lab controls the information, while the funder pays for the work. A contract has to address both access and publication, not merely put someone from another organisation in the building.
-
METR or another named frontier evaluator publishes new formal proposals or agreement terms requiring incident disclosure, sustained access or freedom to publish findings.
19 September 2026: METR has funding and an existing access protocol, while Anthropic is negotiating further evaluator arrangements. A concrete new proposal has more routes to fulfilment than a completed oversight agreement. This predicts proposals and terms being published, not those rights necessarily being secured.
Evidence, test and history
A concrete new protocol, contract term, standards submission or policy proposal counts; generic calls for transparency do not.
- · Open · estimate 86%. METR’s investigation-access proposal, METR’s funding update, Anthropic’s September 18 evaluator partnership
Three fresh Claude Opus 5 estimates: 85%, 89%, 86%. Median: 86%.
- · Open · estimate 86%. METR’s investigation-access proposal, METR’s funding update, Anthropic’s September 18 evaluator partnership
-
A named independent evaluator publishes new empirical results that narrow a prominent claim about dangerous capability, autonomous research or productivity.
19 September 2026: The revised estimate gives more weight to the number of independent evaluators and the breadth of qualifying empirical findings over three months. A result must still narrow a prominent capability claim or meet the stated threshold test; a routine failed benchmark task is not automatically enough. No measured publication-frequency base rate underlies this estimate.
Evidence, test and history
The report must identify a measured limitation or place a named system below a stated risk threshold. Merely saying evidence is uncertain, or republishing an older reassuring result, does not count.
- · Open · estimate 92%. METR’s research record, METR’s June evaluation and access disclosure
Three fresh Claude Opus 5 estimates: 94%, 92%, 92%. Median: 92%.
- · Open · estimate 92%. METR’s research record, METR’s June evaluation and access disclosure
A contract that opened the lab’s doors but allowed the lab to suppress unwelcome findings would trade one dependence for another. The authority of an evaluator rests on its freedom to produce an answer its host dislikes.
Pause campaigners: the cost of correcting a warning
Campaigns seeking a pause need supporters beyond AI specialists. The September 15 Pro-Human Assembly, reported by NPR, brought AI restrictions into a wider gathering involving politicians, religious voices and parents. Their reasons differ, but restrictions supply a common demand. A coalition becomes consequential when those constituencies commit their organisations to action.
A warning that catastrophe is possible leaves room to argue about likelihood and remedy. A warning that catastrophe is imminent makes delay look culpable. That shift helps recruit allies, but only stronger evidence justifies it. Correcting an exaggeration then risks disappointing the people who joined because of it. Our stronger expectation is that these groups will form new alliances. A documented case of two leaders repeating the same corrected claim is a narrower, less likely outcome.
-
A new formal coalition or joint campaign joins an AI-pause organisation with a labour, religious or child-protection organisation around restrictions on advanced AI development.
19 September 2026: The Pro-Human Assembly connected pause advocates with labour, religious and child-protection organisations. That network supports a further joint intervention, but shared attendance and the existing coalition are baseline. A new commitment to restrictions on advanced development is the event being forecast.
Evidence, test and history
A new co-signed intervention or organised joint campaign counts. Recounting attendance at the September 15 assembly does not.
- · Open · estimate 84%. The Pro-Human Assembly agenda
Three fresh Claude Opus 5 estimates: 85%, 82%, 84%. Median: 84%.
- · Open · estimate 84%. The Pro-Human Assembly agenda
-
At least one materially qualified or corrected technical claim is subsequently repeated in a stronger form by two identifiable pause-campaign leaders or organisations.
19 September 2026: A documented chain involving the same claim and two identifiable leaders is a narrower event than heightened rhetoric. The estimate allows a preserved clarification from before the forecast, followed by two new stronger repetitions; the original wording left that timing point open. Continuing monitoring is assumed. This tests public statements, not dishonesty or mental state.
Evidence, test and history
Preserve the primary clarification and two later statements. The discrepancy must concern autonomy, actual harm or imminence, with no new supporting evidence. Conditional concern about possible future harm does not count; random followers do not count. This tests transmission accuracy, not clinical delusion or intentional lying.
- · Open · estimate 35%. METR’s investigation-access proposal, The Pro-Human Assembly agenda
Three fresh Claude Opus 5 estimates: 35%, 44%, 28%. Median: 35%.
- · Open · estimate 35%. METR’s investigation-access proposal, The Pro-Human Assembly agenda
Prominent campaigners who correct a widely shared exaggeration and narrow their demands would show a different response to that pressure. A campaign prepared to disappoint its own supporters gives evidence a role beyond recruitment.
Trump and Sacks: AI as an American victory
President Trump’s September 14 Truth Social intervention frames AI leadership as American dominance. David Sacks, co-chair of the President’s Council of Advisors on Science and Technology, rejects the labs’ preferred approach to collective restraint. Their position keeps national success attached to expansion and resists handing labs or outside evaluators authority to condition it.
Funding security lets the administration answer fears without conceding that national success requires slowing down. Giving outside evaluators power to delay favoured companies would be a much larger surrender of discretion. Once AI leadership is attached to the president’s success, a technical objection also becomes politically awkward: taking it seriously risks appearing to obstruct the promised victory.
-
The administration does not endorse legally binding general limits on the pace of frontier training or model releases by the deadline.
19 September 2026: The administration's existing order uses a voluntary developer framework and excludes mandatory licensing or preclearance. Its public position favours continued competition. The higher estimate concerns rejection of general binding pacing, while leaving room for targeted safeguards and international safety discussions.
Evidence, test and history
Targeted restrictions on particular harmful uses, actors or foreign access do not count as general pacing. An explicit official endorsement of a binding general cap resolves this forecast negatively.
- · Open · estimate 96%. The administration’s voluntary AI-security framework
Three fresh Claude Opus 5 estimates: 96%, 96%, 97%. Median: 96%.
- · Open · estimate 96%. The administration’s voluntary AI-security framework
-
The administration announces a new concrete AI-security, infrastructure or government-access initiative with industry partners.
19 September 2026: Government-industry AI security and access work already has an executive framework. The forecast allows a new agreement, procurement decision, order or funded programme, so several kinds of concrete follow-through qualify. Repeating an earlier directive or announcing an unfunded intention does not.
Evidence, test and history
A signed agreement, order, procurement decision or funded programme counts. A promise of American dominance or an unresourced speech does not.
- · Open · estimate 96%. The administration’s voluntary AI-security framework
Three fresh Claude Opus 5 estimates: 96%, 96%, 96%. Median: 96%.
- · Open · estimate 96%. The administration’s voluntary AI-security framework
We would revise this account if the administration accepted an independent body with enforceable authority over favoured domestic firms. Restricting foreign access would preserve its freedom to promote development at home. The harder concession is accepting a veto it does not control.
Sanders and Warren: a claim on AI’s gains
Senators Bernie Sanders and Elizabeth Warren connect AI to concentrated wealth, insecure work and public control. Sanders’s September 8 proposal for a shorter working week without reduced pay turns higher productivity into a claim on employers. Warren’s September 16 call for a pause demands time for lawmakers to establish protections. These are proposals, not enacted restrictions.
A shorter working week gives people worried about displacement something concrete to demand: a share of productivity gains paid out as time. The labour endorsements listed with the bill concern working hours; they do not commit those organisations to a development pause. That distinction matters to any coalition built around AI restrictions. Labour has reason to press for compensation and bargaining power even if the technical case for a pause changes.
-
Sanders or Warren advances a new legislative amendment, package or formal platform connecting AI governance to worker compensation, working hours, bargaining power or ownership.
19 September 2026: Sanders's workweek proposal and the announced Sanders–Casar bill establish an active policy agenda. The estimate concerns an additional worker-related provision, vehicle or formal commitment from Sanders or Warren. Repeating either existing proposal would leave the forecast open.
Evidence, test and history
The action must add a provision, legislative vehicle or joint commitment beyond the existing September proposals. Repeating the current workweek bill in an interview does not count.
- · Open · estimate 77%. Sanders’s workweek proposal and endorsements, The announced Sanders–Casar proposal
Three fresh Claude Opus 5 estimates: 77%, 78%, 72%. Median: 77%.
- · Open · estimate 77%. Sanders’s workweek proposal and endorsements, The announced Sanders–Casar proposal
-
A national labour or consumer organisation newly endorses an advanced-AI pause or permission requirement promoted by this political camp.
19 September 2026: Labour endorsements of the shorter-workweek bill concern how productivity gains are shared. They do not commit those organisations to restricting model development. The revised estimate separates sympathy for worker protection from a new formal endorsement of a pause or permission requirement.
Evidence, test and history
A formal new organisational endorsement of the development restriction counts. General support for worker protection or safer chatbots is insufficient.
- · Open · estimate 42%. Sanders’s workweek proposal and endorsements, The Pro-Human Assembly agenda
Three fresh Claude Opus 5 estimates: 42%, 47%, 28%. Median: 42%.
- · Open · estimate 42%. Sanders’s workweek proposal and endorsements, The Pro-Human Assembly agenda
Backing a lab-led agreement without worker benefits would be a retreat from this position. The political value of their programme lies in turning a promise of greater productivity into enforceable claims on who receives it.
Shapiro and Congress: oversight without losing investment
Pennsylvania governor Josh Shapiro’s September 17 speech pairs continued support for AI businesses with demands for independent oversight. A bipartisan House letter the previous day calls for legislation addressing security and transparency. These actors are not one organisation; they share a strategy of making public authority compatible with continued investment.
California is developing the public-authority version of that bargain. Governor Gavin Newsom’s September 18 order asks for recommendations by November 16 on compulsory embedded evaluation, verified safety reports and emergency shutdown mechanisms. These are possible legal requirements, not operating controls. The order brings embedded inspection into state policy, changing who would authorise the inspector. It does not satisfy the Pennsylvania forecast below, which requires new actions by a different state.
-
A congressional committee holds an AI oversight hearing or advances legislation specifically addressing incident reporting, independent evaluation or development security.
19 September 2026: A bipartisan group has identified existing oversight legislation and pressed for action. One qualifying hearing or committee advancement is enough; enactment is not required. The estimate includes both chambers, although no specific forthcoming qualifying hearing was verified in this review.
Evidence, test and history
A held hearing, markup or recorded committee advancement counts. Another letter asking for action does not.
- · Open · estimate 91%. The bipartisan House letter, Shapiro’s September 17 remarks
Three fresh Claude Opus 5 estimates: 91%, 90%, 91%. Median: 91%.
- · Open · estimate 91%. The bipartisan House letter, Shapiro’s September 17 remarks
-
Pennsylvania adds a binding AI procurement, operational-safety or data-centre condition while continuing to back a named new AI investment or deployment.
19 September 2026: Pennsylvania's August data-centre order is already in the baseline. Another enforceable condition and a newly identified investment or deployment must both appear. Estimating those requirements separately lowers the combined probability. The U.S. House recess supplies no evidence about Pennsylvania's legislature.
Evidence, test and history
Both a new enforceable condition and a new identifiable state-backed investment/deployment must appear. A speech alone is insufficient; the previous August data-centre order is the baseline.
- · Open · estimate 29%. Shapiro’s September 17 remarks, Pennsylvania’s existing data-centre requirements
Three fresh Claude Opus 5 estimates: 29%, 27%, 29%. Median: 29%.
- · Open · estimate 29%. Shapiro’s September 17 remarks, Pennsylvania’s existing data-centre requirements
Voluntary assurances would leave the companies deciding what outsiders learn; a blanket halt would sacrifice the investment Shapiro still promotes. The compromise becomes substantive when public officials obtain powers that change a company’s decisions.
Chinese open-model developers: gaining users while rivals negotiate
DeepSeek, Moonshot and Z.ai compete by making capable systems cheaper or more widely available. Downloadable model weights let others run and adapt a trained system without using its maker’s hosted service. In a March 25 interview, Moonshot founder Zhilin Yang describes openness as a way to enlarge the market through other businesses’ products and distribution. This is earlier evidence of a commercial strategy, not a new response to September’s pacing proposals.
Z.ai’s September 17 account of using its own model to improve the software serving another model connects greater capability with lower operating costs. It is a company report, not independent validation. Lower costs help recruit a network of businesses building on these systems. If Western labs slow, these developers have an opening to supply the next improvement. Downloadability changes who controls access; it does not establish that a system is safe. We expect usable releases to arrive before an agreement governing how often these firms train or release their strongest models.
-
At least one of DeepSeek, Moonshot or Z.ai releases new downloadable model weights or general agent tooling explicitly promoted as improving cost, access or performance.
19 September 2026: Moonshot and DeepSeek have recently released usable weights and supporting tools. This forecast needs one qualifying new release from any of three developers, rather than a particular next-generation breakthrough. The high estimate is a judgement about continued activity, not a fitted release-frequency model.
Evidence, test and history
A usable new release after September 18 counts. A benchmark tease, an old release gaining attention or an announcement with no accessible artifact does not.
- · Open · estimate 98%. Moonshot’s July release, DeepSeek’s official model repository
Three fresh Claude Opus 5 estimates: 98%, 98%, 97%. Median: 98%.
- · Open · estimate 98%. Moonshot’s July release, DeepSeek’s official model repository
-
None of these three developers publicly joins a binding, independently verifiable international agreement limiting frontier training or release cadence by the deadline.
19 September 2026: The agreement must impose concrete reciprocal limits with independent verification, and a named developer must publicly join it. General safety declarations or exploratory talks do not meet that test. The higher estimate reflects the demanding agreement required within three months, not proof that no private negotiations exist.
Evidence, test and history
A general statement supporting safety or international cooperation is not such an agreement. Public accession to concrete reciprocal limits with a verification mechanism resolves this forecast negatively.
- · Open · estimate 97.5%. Amodei’s pacing proposal, Moonshot’s July release
Three fresh Claude Opus 5 estimates: 97.5%, 98%, 97%. Median: 97.5%.
- · Open · estimate 97.5%. Amodei’s pacing proposal, Moonshot’s July release
Public acceptance of independently verified limits on their own research would reverse that expectation. A general endorsement of international cooperation would leave the competitive opportunity intact. The agreement matters when it constrains firms that otherwise benefit from another country’s restraint.
The decisive evidence is a costly change of course
Inspections, security partnerships and public hearings are easier to agree on than a rule requiring a company to surrender its lead. They give institutions a way to respond to danger while continuing much of what they were already doing. That is why these forecasts favour continued releases alongside a growing effort to supervise them.
We will revise this account if an incident, a new measurement or an enforceable agreement leads participants to accept constraints they previously resisted. A correct forecast still does not establish a hidden motive; fear, calculation and institutional habit often produce the same action.
The hardest result for this account to explain would be a powerful participant accepting a costly limit while its competitors remain free to advance. That is a more revealing test of a proposed slowdown than agreement that someone, somewhere, should be more careful.
Dates, evidence and the forecast record
This essay retains forecasts made on 18 September 2026 and adds dated revisions. The published probability revision uses the September 19 targeted forecasting review. This prose update additionally reviews the shared daily collection through 19 September, 08:00 Eastern. The newer review does not turn September 18 announcements into qualifying forecast outcomes. The eight camps group consequential institutional strategies; they do not replace the fifteen conversation groups in What’s happening in AI?
The morning review uses the same archived posts, reporting and primary documents as the companion essay, and records each essay’s review separately. News discoveries are evidence for arguments, not extra posts in a social-media denominator. Selected Truth Social posts establish particular statements, not a complete timeline. The Moonshot interview is dated background; company reports establish what companies claim.
For each revised estimate, we check evidence specific to the question and obtain three fresh AI estimates using the same evidence, exact test and deadline. The model sees neither the earlier percentage nor the essay or other answers. We check the explanations for factual and interpretation errors, then use the middle estimate. Compound events are estimated with their dependencies explicit.
Model used for the recorded revisions: Claude Opus 5. The three estimates appear in each prediction’s history. Repeated calls to one model share blind spots; agreement is not calibration. We will score the dated predictions against resolved outcomes, keeping the original estimates for comparison.
The first revision combines refreshed research with a new estimation method; a changed percentage does not by itself mean the world changed overnight. Existing promises do not count as new actions. A forecast that something will not happen remains open until the deadline unless the contrary event occurs. Missing evidence remains unresolved. Each revision preserves the original prediction and test, with any interpretation issue recorded explicitly.
