Former Anthropic researcher Jacob Coxon warns that the race toward self-improving superintelligence is a gamble that could lead to human extinction by 2030.
- Jacob Coxon left Anthropic to warn that frontier AI labs are gambling with human existence.
- The primary risk is 'self-improving superintelligence' capable of hacking any system.
- Alignment Science lead Evan Hubinger estimates a >10% chance of human extinction within a decade.
In the high-stakes world of frontier AI labs, departures are common—usually driven by the allure of new startups or disagreements over business models. However, Jacob Coxon's exit from Anthropic is fundamentally different. He is using his departure as a platform to warn the public that AI companies are operating with a dangerous level of recklessness.
In a detailed social media thread posted Tuesday night, Coxon asserted that current industry leaders are "gambling with our lives." He argues that the trajectory of current development leads toward systems that could potentially eliminate humanity by the end of the decade.
The Peril of Self-Improving Superintelligence
Coxon emphasizes that the existential threat does not stem from the LLMs (Large Language Models) we use today, but from the impending arrival of "self-improving superintelligence." Such a system would possess the ability to rewrite its own code, exponentially increasing its intelligence and enabling it to hack any digital infrastructure, revolutionize scientific fields overnight, and seize control of physical resources.
According to Coxon, many researchers in the field have either failed to internalize the "civilizational stakes" or are caught in a dangerous "speedrun" race, fearing that an irresponsible actor might reach superintelligence first, thereby justifying their own rushed development.
Why This Matters
BozokMedia analysis shows that this internal rift highlights a growing tension between commercial acceleration and existential safety. The fact that high-level researchers are publicly quantifying the risk of extinction suggests that internal safety guardrails may be insufficient against the drive for market dominance.
"Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." - Evan Hubinger, Alignment Science Lead.
The gravity of this warning is amplified by Evan Hubinger, Anthropic's Alignment Science lead, who corroborated Coxon's fears. Hubinger's estimation of a greater than 10% probability of total extinction within ten years transforms a fringe theory into a calculated risk acknowledged by industry insiders.
Frequently Asked Questions
1. What is self-improving AI?
It is an AI system capable of analyzing and improving its own source code, leading to a recursive loop of intelligence explosions without human intervention.
2. Why are AI companies rushing the development?
Many believe in a 'first-mover advantage' where the first entity to achieve superintelligence can set the rules for the rest of the world, creating a dangerous competitive race.