Forecast archive
Every original prediction and dated revision, with the probability, resolution test and explanation preserved. New estimates add to this record; they do not replace earlier ones.
All predictions in this round are due by 18 December 2026. Read each history from its original estimate through later revisions and outcomes.
Download dated records
Frontier labs: Anthropic and OpenAI
At least one of Anthropic or OpenAI publicly establishes a named arrangement for embedded external evaluation with concrete access and disclosure terms.
Resolution test: An agreement, published terms or confirmation that access is operating counts. Repeating the September promise does not.
Anthropic has named Accenture, but its September 18 announcement still leaves operating and disclosure standards unsettled. That announcement belongs to the baseline. A new arrangement or confirmation of operating access needs to satisfy the concrete test; another promise alone is insufficient.
Anthropic’s September 18 evaluator partnership, METR’s investigation-access proposal
Three Claude Opus 5 estimates: 63%, 64%, 68%.
Both companies launch or substantially expand paid model or agent products before binding industry-wide pacing restrictions take effect.
Resolution test: A public new release or substantial expansion at each company, after September 18 and before any binding industry-wide pacing regime. This is compatible with their stated proposals and is not itself evidence of hypocrisy.
Anthropic explicitly plans to continue training and releasing models alongside outside evaluation, and OpenAI is expanding its paid offerings. The estimate rises because each company has several routes to a qualifying expansion before a binding industry-wide pacing regime takes effect.
Anthropic’s September 18 evaluator partnership, OpenAI’s September 3 release, The administration’s voluntary AI-security framework
Three Claude Opus 5 estimates: 93%, 92%, 94%.
Expansion businesses: Nvidia and Meta
Nvidia announces a new funded initiative, product or government partnership for AI security, monitoring or defensive use without endorsing a general frontier slowdown.
Resolution test: A new identifiable programme, resource commitment, product or agreement counts. A speech about engineering safety alone does not.
Nvidia's existing CrowdStrike partnership demonstrates a route from safety concerns to defensive products. A new security product, funded programme or government agreement would qualify; the existing partnership cannot count again. The estimate concerns further identifiable activity, not endorsement of a common slowdown.
Nvidia and CrowdStrike’s security announcement, Huang’s September 15 remarks
Three Claude Opus 5 estimates: 90%, 93%, 93%.
Meta expands public access to an agent or advanced model and issues a new public objection to general pacing requirements.
Resolution test: Both a new substantive availability/capability expansion and a new attributable policy statement or lobbying position must occur. Existing Muse availability and the August essay are the baseline.
Meta has announced further Muse expansion, but delivering it and issuing a new objection to general pacing are separate requirements. Estimating the second event conditional on the first lowers the combined probability. Neither the September launch nor the earlier policy essay counts as a new action.
Meta’s Muse launch and plans, Zuckerberg’s earlier essay
Three Claude Opus 5 estimates: 70%, 74%, 72%.
Independent evaluators, including METR
METR or another named frontier evaluator publishes new formal proposals or agreement terms requiring incident disclosure, sustained access or freedom to publish findings.
Resolution test: A concrete new protocol, contract term, standards submission or policy proposal counts; generic calls for transparency do not.
METR has funding and an existing access protocol, while Anthropic is negotiating further evaluator arrangements. A concrete new proposal has more routes to fulfilment than a completed oversight agreement. This predicts proposals and terms being published, not those rights necessarily being secured.
METR’s investigation-access proposal, METR’s funding update, Anthropic’s September 18 evaluator partnership
Three Claude Opus 5 estimates: 85%, 89%, 86%.
A named independent evaluator publishes new empirical results that narrow a prominent claim about dangerous capability, autonomous research or productivity.
Resolution test: The report must identify a measured limitation or place a named system below a stated risk threshold. Merely saying evidence is uncertain, or republishing an older reassuring result, does not count.
The revised estimate gives more weight to the number of independent evaluators and the breadth of qualifying empirical findings over three months. A result must still narrow a prominent capability claim or meet the stated threshold test; a routine failed benchmark task is not automatically enough. No measured publication-frequency base rate underlies this estimate.
METR’s research record, METR’s June evaluation and access disclosure
Three Claude Opus 5 estimates: 94%, 92%, 92%.
Pause campaigners and organised safety advocacy
A new formal coalition or joint campaign joins an AI-pause organisation with a labour, religious or child-protection organisation around restrictions on advanced AI development.
Resolution test: A new co-signed intervention or organised joint campaign counts. Recounting attendance at the September 15 assembly does not.
The Pro-Human Assembly connected pause advocates with labour, religious and child-protection organisations. That network supports a further joint intervention, but shared attendance and the existing coalition are baseline. A new commitment to restrictions on advanced development is the event being forecast.
Three Claude Opus 5 estimates: 85%, 82%, 84%.
At least one materially qualified or corrected technical claim is subsequently repeated in a stronger form by two identifiable pause-campaign leaders or organisations.
Resolution test: Preserve the primary clarification and two later statements. The discrepancy must concern autonomy, actual harm or imminence, with no new supporting evidence. Conditional concern about possible future harm does not count; random followers do not count. This tests transmission accuracy, not clinical delusion or intentional lying.
A documented chain involving the same claim and two identifiable leaders is a narrower event than heightened rhetoric. The estimate allows a preserved clarification from before the forecast, followed by two new stronger repetitions; the original wording left that timing point open. Continuing monitoring is assumed. This tests public statements, not dishonesty or mental state.
METR’s investigation-access proposal, The Pro-Human Assembly agenda
Three Claude Opus 5 estimates: 35%, 44%, 28%.
Trump, Sacks and the administration's acceleration position
The administration does not endorse legally binding general limits on the pace of frontier training or model releases by the deadline.
Resolution test: Targeted restrictions on particular harmful uses, actors or foreign access do not count as general pacing. An explicit official endorsement of a binding general cap resolves this forecast negatively.
The administration's existing order uses a voluntary developer framework and excludes mandatory licensing or preclearance. Its public position favours continued competition. The higher estimate concerns rejection of general binding pacing, while leaving room for targeted safeguards and international safety discussions.
The administration’s voluntary AI-security framework
Three Claude Opus 5 estimates: 96%, 96%, 97%.
The administration announces a new concrete AI-security, infrastructure or government-access initiative with industry partners.
Resolution test: A signed agreement, order, procurement decision or funded programme counts. A promise of American dominance or an unresourced speech does not.
Government-industry AI security and access work already has an executive framework. The forecast allows a new agreement, procurement decision, order or funded programme, so several kinds of concrete follow-through qualify. Repeating an earlier directive or announcing an unfunded intention does not.
The administration’s voluntary AI-security framework
Three Claude Opus 5 estimates: 96%, 96%, 96%.
Economic populists: Sanders and Warren
Sanders or Warren advances a new legislative amendment, package or formal platform connecting AI governance to worker compensation, working hours, bargaining power or ownership.
Resolution test: The action must add a provision, legislative vehicle or joint commitment beyond the existing September proposals. Repeating the current workweek bill in an interview does not count.
Sanders's workweek proposal and the announced Sanders–Casar bill establish an active policy agenda. The estimate concerns an additional worker-related provision, vehicle or formal commitment from Sanders or Warren. Repeating either existing proposal would leave the forecast open.
Sanders’s workweek proposal and endorsements, The announced Sanders–Casar proposal
Three Claude Opus 5 estimates: 77%, 78%, 72%.
A national labour or consumer organisation newly endorses an advanced-AI pause or permission requirement promoted by this political camp.
Resolution test: A formal new organisational endorsement of the development restriction counts. General support for worker protection or safer chatbots is insufficient.
Labour endorsements of the shorter-workweek bill concern how productivity gains are shared. They do not commit those organisations to restricting model development. The revised estimate separates sympathy for worker protection from a new formal endorsement of a pause or permission requirement.
Sanders’s workweek proposal and endorsements, The Pro-Human Assembly agenda
Three Claude Opus 5 estimates: 42%, 47%, 28%.
Pragmatic oversight: Shapiro and bipartisan legislators
A congressional committee holds an AI oversight hearing or advances legislation specifically addressing incident reporting, independent evaluation or development security.
Resolution test: A held hearing, markup or recorded committee advancement counts. Another letter asking for action does not.
A bipartisan group has identified existing oversight legislation and pressed for action. One qualifying hearing or committee advancement is enough; enactment is not required. The estimate includes both chambers, although no specific forthcoming qualifying hearing was verified in this review.
The bipartisan House letter, Shapiro’s September 17 remarks
Three Claude Opus 5 estimates: 91%, 90%, 91%.
Pennsylvania adds a binding AI procurement, operational-safety or data-centre condition while continuing to back a named new AI investment or deployment.
Resolution test: Both a new enforceable condition and a new identifiable state-backed investment/deployment must appear. A speech alone is insufficient; the previous August data-centre order is the baseline.
Pennsylvania's August data-centre order is already in the baseline. Another enforceable condition and a newly identified investment or deployment must both appear. Estimating those requirements separately lowers the combined probability. The U.S. House recess supplies no evidence about Pennsylvania's legislature.
Shapiro’s September 17 remarks, Pennsylvania’s existing data-centre requirements
Three Claude Opus 5 estimates: 29%, 27%, 29%.
Chinese open-model developers: DeepSeek, Moonshot and Z.ai
At least one of DeepSeek, Moonshot or Z.ai releases new downloadable model weights or general agent tooling explicitly promoted as improving cost, access or performance.
Resolution test: A usable new release after September 18 counts. A benchmark tease, an old release gaining attention or an announcement with no accessible artifact does not.
Moonshot and DeepSeek have recently released usable weights and supporting tools. This forecast needs one qualifying new release from any of three developers, rather than a particular next-generation breakthrough. The high estimate is a judgement about continued activity, not a fitted release-frequency model.
Moonshot’s July release, DeepSeek’s official model repository
Three Claude Opus 5 estimates: 98%, 98%, 97%.
None of these three developers publicly joins a binding, independently verifiable international agreement limiting frontier training or release cadence by the deadline.
Resolution test: A general statement supporting safety or international cooperation is not such an agreement. Public accession to concrete reciprocal limits with a verification mechanism resolves this forecast negatively.
The agreement must impose concrete reciprocal limits with independent verification, and a named developer must publicly join it. General safety declarations or exploratory talks do not meet that test. The higher estimate reflects the demanding agreement required within three months, not proof that no private negotiations exist.
Amodei’s pacing proposal, Moonshot’s July release
Three Claude Opus 5 estimates: 97.5%, 98%, 97%.
