Anthropic’s head of alignment science stated publicly that the probability of AI killing all humans is greater than 10% within the next decade. Is the number real? Rather than adjudicating the figure, this article reads the “structure of the competition” the statement exposed and the difficulty it creates, then presents the realistic options available under uncertain information as implications for Japan.
A Startling Message From Inside Frontier AI
On 8 September 2026, Anthropic researcher Jacob Coxon announced his resignation on X (formerly Twitter). He had spent roughly three years at OpenAI and Anthropic researching the pretraining of AI models. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither company is acting responsibly, that they are racing straight to self-improving superintelligence and gambling with our lives. He forfeited his equity and left the industry. His post had passed 150 million views as of 10 September (Fortune).
Eighty-two minutes after Coxon’s post, Evan Hubinger, Anthropic’s head of alignment science, replied as follows.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Hubinger posted the following clarification shortly afterwards.
To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement.
The statements by Coxon and then Hubinger have sent a considerable shock through the American AI industry and its politics. Within days, calls for AI regulation in Congress intensified sharply. The AI Kill Switch Act, introduced in July by Representatives Moran and Lieu to require companies to retain the means to halt or suspend their AI systems, drew renewed attention, and in the Senate, Senators Cruz, Thune and Klobuchar are negotiating legislation aimed at catastrophic risks. Senator Klobuchar said the bill should include “requiring developers to work with government experts to verify and test models to make sure AI is safe.” At the state level, on 9 September, Governor Newsom of California signed two bills creating independent verification organisations for AI and a registry for AI auditors.
To organize the argument a little: the core of this whole debate is recursive self-improvement, or RSI. It refers to the process in which AI repeatedly improves itself and evolves without depending on humans. Once that process takes hold, AI could advance far more dramatically than it does today.
On recursive self-improvement, as it happens, Jakub Pachocki, OpenAI’s Chief Scientist, had addressed the subject two days before Coxon’s post, in an essay titled “An Alien Mind.” He expects that if AI progress continues, recursive self-improvement will be at the very core of future scientific discovery, and he states that for OpenAI it is the only way to remain at the frontier of AI research. On the risks that come with it, however, he writes that “this is a time that calls for extreme caution.” He names two levers for addressing the risk. One is to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop. The other is to coordinate to slow down development as far as needed to build confidence in those measures. The best way forward he currently sees is a combination of the two.
Then, on 12 September, Anthropic CEO Dario Amodei published an essay titled “We Must Pace the Frontier,” writing that “we must slow the pace at which we improve the capabilities of AI models,” and proposing a framework for slowing down in three stages: each frontier company accepting an embedded team of external third-party evaluators; frontier companies within democratic countries coordinating on common safety standards and limits on the rate of progress; and democratic governments coordinating with authoritarian governments to the extent that this is possible. OpenAI CEO Sam Altman responded that “I agree with Dario that we need to pace the frontier.”
What Game Theory Reveals About Why Cooperation in the AI Race Is Hard
On the point that recursive self-improvement in AI carries serious risk, Jacob Coxon, Evan Hubinger, Jakub Pachocki and Dario Amodei are in agreement. Their responses differ. Coxon argues it should be stopped. Hubinger puts a time-bounded probability and a magnitude on the risk. Pachocki discusses the possibility of solutions, and Amodei proposes concrete measures. That Sam Altman endorsed the need to pace the frontier further indicates how widely this recognition is shared. In other words, people working inside frontier AI hold the same view of the risk, yet their responses diverge — and the gap looks widest between Coxon, who wants a halt, and Pachocki and Amodei, who want to build a framework while continuing to advance.
Let us gather the elements of this argument into one simple expression. The players in frontier AI face the following equation.
Expected value of competing in AI =
probability of winning × reward for winning − probability of losing × penalty for losing − probability of AI risk × magnitude of AI risk
Consider a frontier AI company such as Anthropic or OpenAI. If it survives the race, it gains not only an enormous share of the market and the revenue that comes with it, but also a continuing supply of abundant new investment. If it loses, it forfeits future opportunity and puts the business base it has built so far at risk as well.
This is not only a corporate-level matter. Bringing the US–China AI competition into view reveals the asymmetry of outcomes that the equation implies. For the United States, which currently leads the AI industry worldwide, continuing to win preserves its present competitive advantage, whereas losing could set off large chain effects not only in AI but across the economy, politics, the military and more. That is why US Treasury Secretary Scott Bessent said, in a speech in Washington on 8 September 2026, that “there is no day after tomorrow if China wins at this,” and that if China “were to pull away from us on AI, then nothing else would matter.” The magnitude of the penalty for losing, in short, far exceeds the reward for winning.
Seen from China, the situation is roughly reversed: the reward for winning is overwhelmingly larger than the penalty for losing. The win-or-lose outcome represented by the first and second terms is zero-sum among the players. One player’s win is another’s loss.
Viewed this way, it is easy to see why the third term, the AI-risk term, tends to be deferred in one decision after another. Unlike the reward of the first term and the penalty of the second, the third term is borne not by the player alone but by society as a whole and ultimately by the entire world. The third term is, in effect, an external term: any single player bears only a fraction of its cost, and no single player can control it either.
Because it is external, the third term also carries greater uncertainty than the first and second. Unless the information sharing and cooperation needed across corporate boundaries, and across national ones, are in place, even quantifying its probability and the scale of the harm is extremely difficult.
It is certainly conceivable that companies and countries could cooperate to slow frontier AI development and thereby reduce the risk in the third term, as Jakub Pachocki and Dario Amodei propose. But unless trust has been established among the players, slowing down on one side alone merely lowers one’s own probability of winning, and a player reasoning rationally will not take that action. Harder still: between one company and another, a state or a public institution can impose discipline, yet states themselves have become AI players, and at present there is no arbiter to discipline states over the development of AI.
Implications for Japan: Finding the Best Realistic Answer Under Uncertain Information
How, then, does Japan relate to this debate? At its core, this AI-risk debate is about the risk created by recursive self-improvement in frontier AI, the possibility of defending against that risk, and the management of development pace needed to do so. As the equation above shows, the win-or-lose outcome of the first and second terms is a matter among frontier AI players, and Japan is not directly affected by it. Even so, although Japan is not a frontier AI player, it cannot escape the economic and geopolitical chain effects that the outcome sets off.
As for the externality of AI risk shown in the third term, Japan bears that risk just as every other country does, and is therefore a direct stakeholder.
Seen in that light, and despite how much about AI risk remains uncertain, the shape of the realistic course for Japan comes into view. First, the direction of raising corporate productivity and profitability through the use of today’s AI is largely unrelated to the risks of frontier AI development. Second, raising Japan’s own AI capability does not, until it stands shoulder to shoulder with the frontier players, meaningfully push up the risk of the third term, so it gives no reason to slow down. If anything, Japan can treat the current leading edge of frontier AI as its safety line and, up to that line, apply and advance AI at full speed.
On the third term, Japan may be able to contribute by helping build an environment for dialogue and a basis of trust among the frontier AI players, and through such activity help reduce AI risk. Precisely because Japan is not a frontier AI player, it holds a position from which it can make proposals as a third party.
The figure quoted at the opening of this article, the “greater than 10%” offered by Anthropic’s head of alignment science Evan Hubinger, is not something that can be verified. What this debate over frontier AI risk has exposed, however, is another matter: the asymmetry between reward and penalty in frontier AI competition, the externality of AI risk, the dual-player structure of companies and states, and the present absence of any arbiter to discipline states. Together these greatly increase both the complexity of the problem and the uncertainty of how it will unfold. For Japan, though, reading that structure is comparatively simple. Since Japan is not a party to the frontier race, it has no reason to slow down — up to the point where it stands shoulder to shoulder with the leading edge, it should apply and develop AI at full speed. At the same time, precisely because it is not a party, it can become a place from which dialogue is proposed. How to make use of those two strategies, and how to generate the greatest effect from them, is the real question for Japan.
References and Sources
- TIME — “He Helped Build Powerful AI at OpenAI and Anthropic. Now He’s Afraid It Could Kill Us” — Content of Jacob Coxon’s resignation post, his age, and his tenure at both companies (direct access outside a browser may be restricted)
- Fortune — “Ex-Anthropic researcher Jacob Coxon says AI could end humanity…” — More than 150 million views as of 10 September
- Evan Hubinger — post on X (8 September 2026) — “>10% within the next decade,” and the statement that there is no plan to solve alignment for superintelligence
- Evan Hubinger — follow-up post on X (8 September 2026) — Clarification that risk from present models is low and that the concern is superintelligence arising from recursive self-improvement
- Anthropic — “Risk Report, August 2026” (PDF) — The company’s own risk assessment, referenced by Hubinger in his follow-up post
- Jakub Pachocki (OpenAI) — “An Alien Mind” (6 September 2026) — The outlook for recursive self-improvement, “extreme caution,” and the two levers and their combination (direct access outside a browser may be restricted)
- Dario Amodei (Anthropic) — “We Must Pace the Frontier” (12 September 2026) — The argument for slowing the pace of capability improvement, and the three-stage framework for doing so
- NBC News — “Sam Altman backs Anthropic CEO’s call to slow down the global AI race” (12 September 2026) — Altman’s statement on X agreeing with Amodei
- Al Jazeera — “US legislators push AI safety laws amid human extinction warnings” (11 September 2026) — Overview of AI regulatory legislation moving in the US Congress
- Reuters (via The Spokesman-Review) — “U.S. Senate negotiators consider requiring AI firms to mitigate known major risks” — Legislative negotiations among Senators Cruz, Thune and Klobuchar, and Klobuchar’s statement
- Rep. Ted Lieu — “Reps Lieu and Moran Introduce Bill to Require Kill Switch for AI Systems” (23 July 2026) — Introduction and contents of the AI Kill Switch Act
- Office of the Governor of California — “Governor Newsom signs first-in-the-nation AI safeguards…” (9 September 2026) — Signing of SB 813 and AB 1405, creating independent verification organisations and a registry for AI auditors
- Bloomberg (via Yahoo Finance) — “Bessent Warns Nothing Would Matter If China Wins the AI Race” — Remarks by US Treasury Secretary Scott Bessent on 8 September 2026