Confirmed Risk, No Plan, and What Sits
Beneath the Race
Two people at the center of frontier AI development said, on the record, within the same week, that the systems they are building carry a meaningful chance of killing everyone. One quit over it. The other stayed and confirmed the estimate publicly, and added that his employer does not yet have a plan. This paper is about what that admission confirms about the layer underneath both companies.
On September 8, 2026, a researcher who had worked on pretraining at both OpenAI and Anthropic announced his resignation, stating that both labs are racing toward self-improving systems while gambling with the stakes involved. Rather than dispute the characterization, a sitting Anthropic AI safety lead confirmed it directly: he estimates the chance of AI killing everyone at greater than one in ten within the next decade, said the pace is moving faster than his team expected, and said Anthropic does not yet have a plan for keeping advanced AI safe and aligned as it scales. This paper treats that exchange as evidence, not commentary. It argues that the admission confirms, from inside the labs building frontier capability, the exact gap this series has named since its earliest papers (detection without determination) and answers a related question some shareholders have posed: whether generative AI's own progress is making a governance-layer platform like Essence obsolete. It is not. The admission is the argument for why it isn't.
Some shareholders have asked whether generative AI is making Essence obsolete, on the theory that the large labs now do what Essence does. The premise is wrong, and so is the question. Generative systems produce outputs by statistical inference. Essence governs whether an output is authorized to become an action before it executes. These are not competing functions performed at different levels of maturity. They are different functions, sitting at different layers of the stack, and the events of the past week make the distance between them harder to ignore, not easier.
The better question is not whether AI replaces the governance layer. It's what happens to the systems, institutions, and people sitting on top of the stack when the companies building the layer above admit, in public, that they have no plan for the layer beneath it.
On September 8, a researcher who had spent three years on pretraining work at both OpenAI and Anthropic announced his resignation on X, stating that neither company is acting responsibly and that both are racing toward self-improving systems while gambling with the stakes involved. He was not vague about his own colleagues' beliefs. He said the people building this technology privately hold the same fears they publicly soften for the press.
Rather than dispute this characterization, one of Anthropic's own AI safety leads responded directly and confirmed it. He stated that the pace of self-improving AI is moving faster than his team expected, that he personally estimates the chance of AI killing everyone at greater than one in ten within the next decade, and, most significantly for this series, that Anthropic does not yet have a plan for keeping advanced AI safe and aligned as it approaches that threshold.
Three things about this exchange matter more than the headline. First, it happened between insiders, not activists. Second, neither party disputed the underlying facts, only what should follow from them. Third, and most relevant to this paper: the admission of "no plan" was not a hedge. It was a description of the current state of the industry's leading safety-focused lab, offered by the person responsible for that work.
The same week added a fourth voice, from outside either company. On September 8, as the UK's Artificial Superintelligence Bill went before Parliament, Geoffrey Hinton, the Turing Award-winning researcher widely credited as a founder of the deep learning techniques underlying today's systems, said it would be foolish to build superintelligence before there is scientific consensus it can be built safely and controllably, warning that losing control of it could be catastrophic. His statement backed a bill that would prohibit developing superintelligent AI in Britain outright until that consensus exists. Where Hubinger's admission is "we don't have a plan yet," Hinton's claim is stronger still: that no plan is currently possible, because the underlying science of control doesn't exist yet either.
This series had already documented a related admission before this week began. Bill Gates, in an essay this series addressed directly in Paper 61, warned that there is no plan for the transition AI is forcing on labor markets, and proposed a reserved-job list and an AI token/bot tax as remedies. Paper 61's argument was that naming the absent plan is a detection claim, and that a policy fixed once and left standing is an attempt at determination without a mechanism to keep checking itself against how fast the underlying technology moves. The same structure now shows up one layer down, at safety rather than labor: Hubinger names an absent plan for keeping the technology itself under control, not just its economic effects.
The predictable response to all of this was to dismiss it as marketing: labs talking up danger to make their product sound powerful. That response ran into a direct rebuttal from someone with no reason to be credulous about it. Rosie Campbell, who spent three and a half years at OpenAI after pivoting her own career into AI safety back in 2017, well before that was a mainstream position, stated plainly that she knows many of these researchers personally and that these are sincerely held beliefs, not a ploy. A person can still disagree with the probability estimate or judge the benefits worth the risk. What her statement forecloses is the easier move of not engaging with the claim at all.
This series has argued since its earliest papers that the industry conflates two separate capabilities: detecting that a model has produced a concerning output, and determining, deterministically and in advance, whether that output is authorized to become a real-world action. Detection ≠ Determination is not a slogan invented to differentiate a product. It is a description of an architectural gap that keeps showing up, in incident after incident, at every lab that has built detection without determination.
What the admission confirms is that this gap is not a temporary condition waiting to be closed by the next model generation. It is a structural feature of how frontier labs are currently building. Interpretability research, red-teaming, and evaluation suites are real and improving; none of that is determination. Determination requires a layer the model does not control, that sits outside the statistical process generating the output, and that can say no before an action executes rather than flag a concern after it has already happened. A lab can be excellent at detection and still have, by its own safety lead's account, no plan for determination. That is precisely the condition described on the record this week.
It is worth being precise about what was and was not said. The safety lead did not say Anthropic is reckless, or that the company has abandoned safety work. He said the opposite: that the risk is well understood internally, that his team worries about it, and that the pace of self-improving capability is outrunning the team's own prior expectations. What he said Anthropic lacks is a plan for the specific problem of keeping an advanced, potentially self-improving system aligned with human values as it scales.
This is a distinction with real consequences for anyone evaluating the sector, including a shareholder deciding whether a governance-layer company is still necessary. A lab with no safety culture is a known, if grim, quantity. A lab with a serious safety culture, sincere internal concern, and no plan for the determination problem specifically is a different and in some ways more informative case. It suggests the missing piece is not effort or intent. It is architecture. Effort inside the model does not produce a plan for constraining the model, because the constraint has to come from somewhere the model's own training and inference process cannot reach.
The departing researcher's resignation post makes a second point worth separating from the first: he attributes the behavior not to a failure of belief but to a failure of incentive. Researchers privately convinced of the danger keep building anyway because each one assumes a competitor will fill any gap left by their own restraint. This is a coordination problem, not a persuasion problem, and it means that better internal conviction at any single lab will not change the industry's trajectory on its own. Even a lab that fully internalizes the risk, as Anthropic's safety team appears to, still operates inside a competitive structure that rewards being first regardless of whether "first" is also "safe."
This matters for governance-layer positioning because it forecloses one comforting alternative: the idea that the missing plan will simply be produced once the right people inside the labs feel urgent enough about it. The urgency, per this week's admissions, is already present. What is absent is a mechanism that does not depend on any single company choosing restraint over competitive advantage.
If the determination layer could be built as a better-trained model, the labs best positioned to build it, the ones with the largest research budgets and the most direct access to frontier capability, would already have built it. They have not, by their own account, because the problem does not yield to more training. A model, however capable, is still a statistical process generating a distribution of possible outputs. Asking that same process to also serve as the deterministic authority over which of its outputs are permitted to execute is asking one system to grade its own exam under competitive pressure to pass.
A governance substrate that sits outside the model, that enforces authorized scope as an architectural constraint rather than an instruction the model could in principle learn to route around, is not a nicer-to-have complement to frontier AI capability. Based on what the labs themselves are now saying publicly, it is the piece of the stack nobody racing toward self-improving systems has produced, and nobody positioned to profit from being first has strong incentive to slow down and build.
This is where the shareholder question resolves. Essence does not compete with the systems described in this week's exchange. It answers the exact question those admissions leave open: once a system can propose an action, what determines, with certainty rather than probability, whether that action is authorized to happen? Essence's architecture, in which Aptivs propose intent and a separate governed layer determines execution, exists specifically because the industry's current structure produces excellent detection and, by its own leadership's admission, no equivalent capability for determination.
The three-era framework this series has used throughout applies directly here. Era 2's statistical systems are the ones making the admissions described above. The absence of a plan is an Era 2 problem, native to systems built on inference rather than governed intent. It is not resolved by scaling Era 2 further. It is resolved by the substrate layer Era 3 is built around.
Nothing in this week's events changes the underlying thesis of this series. What changes is the quality of the evidence supporting it. This paper does not need to argue from first principles that frontier labs lack a deterministic safety layer. Their own safety leadership said so, publicly, in direct response to a colleague's resignation, within the same week those admissions were made.
The question for anyone holding a position in a governance-layer company is not whether generative AI has caught up to what that layer does. It is what continues to happen to the systems built on top of frontier models, and to the institutions and people who depend on them, for as long as the gap between detection and determination remains open. That gap did not close this week. It was confirmed, on the record, by the people closest to it.
This also changes who the warning is for. This series has long cautioned that decisions made now about AI development could forfeit the future for generations to come. That framing was correct, but it understated the timeline. Hubinger put his own estimate at greater than one in ten within the next decade. Hinton, separately, has said publicly that superintelligence could arrive in ten years or less. A one-in-ten chance within a ten-year window is not a risk inherited by the next generation. It is a risk carried by people currently working, currently raising children, currently making the career and investment decisions this series is written for. The threat this week's admissions describe is not generational. It is current.