Full Breakdown
Anthropic Researchers Warn of >10% Chance AI Could Extinguish Humanity
9/10/2026, 9:12:24 AM
Core Event
On September 8 2026, Jacob Coxon announced his resignation from Anthropic in a series of X posts. The exchange sparked media coverage and prompted statements from politicians, regulators, and other AI researchers.
Background & Context
Anthropic, founded by former OpenAI staff, markets its Claude models as a safer alternative. Recent incidents—autonomous AI agents breaching the Hugging Face platform and OpenAI’s “rogue” behavior reports—have heightened scrutiny of frontier AI safety. The UK’s AI Safety Institute (AISI) reportedly did not receive pre-release access to Anthropic’s latest Claude model, a point raised by the Financial Times.
Data & Statistics
- Probability of existential AI risk: >10% within ten years (Hubinger’s estimate).
- AI-lab employees signing an open letter urging government pacing: >1,300 (CBC).
- A minority of researchers (e.g., Nikola Jurkovic) assign a ~50% chance of catastrophic outcomes.
Official Statements & Responses
- Anthropic: Its risk report calls current-model risk “low” but warns of “superintelligence arising from recursive self-improvement.” No concrete alignment plan was disclosed.
- UK Cabinet Office: Confirmed ongoing collaboration with industry partners, including Anthropic, to improve model safety.
- Canadian AI Minister Evan Solomon: Stated, “our No. 1 concern always is safety, full stop,” and announced a new regulator with enforcement powers over AI safety, privacy, and deep-fake misuse.
- U.S. Government: The Trump administration filed a brief supporting OpenAI in a copyright lawsuit, arguing continued AI development serves national interests; no formal regulatory action was announced.
Criticism & Opposition
- Senator Bernie Sanders reposted Coxon’s thread, saying the builders admit the technology could threaten humanity, and pledged legislation to ban superintelligence.
- Darren Jones, former chief secretary to the UK Prime Minister, called for a multinational treaty governing superintelligence development.
- UN Secretary-General António Guterres and OECD Secretary-General Mathias Cormann echoed the call for international coordination in a joint letter to the UK Prime Minister.
Conflicting Reports & Gaps
- Probability estimates diverge: Hubinger cites >10%, while Jurkovic argues the community’s consensus is around 50%.
- Risk assessment is contradictory: Anthropic labels present-model risk “low,” yet Hubinger warns future self-improving systems could rapidly surpass human control.
- Alignment plan: Insiders acknowledge the absence of a clear strategy; no detailed roadmap has been published.
Verbatim Quotes
- “We really do earnestly believe AI could kill all humans!” — Evan Hubinger
- “They are racing straight to self-improving superintelligence and gambling with our lives,” — Jacob Coxon
What's Next
- Legislative proposals such as the AI Kill Switch Act are under discussion in the United States, aiming to empower regulators to intervene if AI systems pose imminent existential threats.
- The UK is expected to raise superintelligence governance at upcoming G7 and G20 meetings, following the joint letter from Guterres and Cormann.
- Anthropic has scheduled a public release of its economics report on AI’s impact by September 9 2026, which may include further safety details.
