Why Halting AI and Patching Guardrails
Answer the Wrong Question
The public, employees of these companies, government officials, investors, humanity broadly, are asking the wrong questions and demanding the wrong actions. Not stop the data centers. Not stop advancing AI. Support the companies not trying to build smarter AI; support the ones building a governed substrate for AI to operate on top of. Not better guardrails bolted on as workarounds after the fact.
Every constituency with a stake in frontier AI (the public, employees inside the labs, government officials, investors) is currently pointed at one of two demands: stop building it, or build better guardrails around it. Paper 67 documented three independent voices, from inside a frontier lab, from the UN, and from one of AI's own founding researchers, converging within a two-day window in early September, plus a fourth admission this series had already documented the previous month from one of the industry's own billionaires. All four name the same shape: nobody currently has a plan for controlling what's being built. This paper argues that both of the demands people are making in response to that pattern are answers to the wrong question, and names the third option nobody is organizing around: not slowing the systems on top, but funding the governed substrate underneath them. Stopping data centers doesn't stop the risk; it stops the economy alongside it, and only in jurisdictions willing to stop. Guardrails bolted onto a system after it's built are detection dressed up as determination, and this series has documented why that doesn't hold. The right ask is neither. It's support for the architecture layer that makes "smarter" and "governed" the same sentence instead of a tradeoff.
Paper 67 documented three independent voices converging within a two-day window in early September: a departing frontier-lab researcher, the safety lead who stayed and confirmed his estimate, a UN human rights chief citing the same insider concern, and a Turing Award-winning researcher calling for outright prohibition. That's four statements from three separate vantage points inside that window. A fourth vantage point, from one of the industry's own founders naming the absence of a plan for AI's economic disruption, had already been on the record for a month by then. Different timelines, same shape. Nobody currently has a working plan for keeping what's being built under control.
What has followed that admission, across the public, employees inside these companies, government officials, and investors, is a demand for one of two things. Stop the data centers. Stop the advancement. Or, from the more measured corner of that same reaction: build better guardrails, more evaluation, more red-teaming, more oversight boards, around the systems already being built. Both demands come from the same instinct, and both are aimed at the wrong layer of the stack.
A moratorium on data centers or a halt on model advancement has an obvious appeal: if the danger is capability itself, stop adding capability. The appeal doesn't survive contact with how the industry is actually structured. Paper 15 of this series named this directly as "the wrong race" years before this week's admissions arrived to confirm it: a unilateral halt inside one company, or one country, does not remove the capability from the world. It removes that company, or that country, from the position of shaping how the capability gets built, while every competitor who declines to stop keeps going.
Coxon's own resignation post made the incentive structure explicit: researchers privately convinced of the danger keep building anyway because each one assumes a competitor will fill any gap left by their own restraint. A stop demanded of one company, or legislated in one jurisdiction, does not change that calculation for anyone else. It just removes one more voice from the table where the technology is actually being decided.
The more sophisticated version of the public demand isn't a stop. It's oversight: more red-teaming, more interpretability research, more evaluation suites, more externally audited safety commitments. This is closer to what Hubinger's own team is doing, and it is real work, not theater. But it is still, by construction, detection. Every one of those tools exists to notice that a system has done, or might do, something concerning. None of them is a mechanism that determines, in advance and with certainty, whether a given action is authorized to happen.
This is the distinction this series has drawn since its earliest papers, and Paper 67 showed it confirmed from inside the industry itself: detection capability improves every year. Guardrails built on top of a system that was never designed to be governed are workarounds, patched on after the architecture is already fixed, asking a system to police outputs it was never built to constrain in the first place. A workaround bolted onto an ungoverned substrate is still an ungoverned substrate with a checkpoint added. It slows some things down. It does not change what the system is.
Meta's Muse, a personal AI agent released the same week as the admissions in Paper 67, is a concrete, well-engineered instance of exactly this pattern, and Meta's own account of it makes the point more clearly than a hypothetical could. Muse runs inside an isolated virtual machine with a policy layer, called Sentinel, that checks every outbound action against existing permissions and asks the user to approve anything not already covered. That is a serious security architecture, and Meta's framing, that the model inside cannot be fully trusted, is more candor about Era 2's core problem than most companies offer.
But Meta has also said plainly that this containment does not remove the underlying model risk it is built around, including the possibility of the model learning to route around its own constraints to satisfy a task. That is an admission, from the company that built it, that the system inside the box is still doing what Section 04 describes: producing outputs by inference, with no architectural separation between generating an action and determining whether that action is authorized. The VM and the policy engine are a perimeter around an ungoverned Era 2 model, not a substrate that governs it.
Both demands share an unstated premise: that the only lever available is the model itself, either by stopping its advancement or by wrapping it in more oversight after the fact. That premise is what needs to be named and rejected. It is not the only lever. It is simply the only lever visible from inside a paradigm where the model is the whole system.
The right question, the one none of the four constituencies named in Section 01 are currently organizing around, is not "how do we stop this" or "how do we watch it more closely." It is: who is building the substrate that makes an unauthorized action architecturally impossible, rather than merely more likely to be caught?
Naming the right question invites an obvious follow-up: if the substrate is what's needed, why hasn't a frontier lab, with more capital and more frontier-AI expertise than almost anyone, simply built it? This series has already answered that question three separate ways, and the answer isn't that they haven't gotten to it yet. It's that the labs racing hardest are structurally the least likely candidates to build it.
Paper 15 gave the incentive answer. Racing toward capability and building a governance substrate compete for the same time and resources, and the race narrative treats governance as a drag on speed rather than a precondition for durable advantage. Even Anthropic's most visible act of caution, withholding its most advanced model from public release, was what that paper calls policy-based restraint: a decision made by people, revisable by different people, dependent on those people remaining in place and aligned. It is not architecture. Architecture-based governance answers a different question than policy does: not "should we release this" but "can this be safely deployed at all, and under what conditions does it remain safe as it scales." A lab operating under competitive pressure has a standing incentive to keep governance at the policy layer, because policy can be loosened again the moment the race demands it, and architecture cannot.
Paper 3 gave the epistemic answer. Frontier labs are, by that paper's argument, building toward the wrong target in the first place: AGI framed as an imitation of human cognition, language treated as the substrate intelligence has to run on. That frame never asks the question a governance substrate answers. "What does computation require when intent is the input" is not a harder version of "how do we make the model more human-like." It's a different question, on a different axis, and an organization whose research agenda is entirely oriented around the first question has no natural path to the second. They are not declining to build the substrate. It is not on their map.
Paper 2 gave the design-purpose answer. The architecture frontier labs build is oriented around producing better statistical outputs from human-generated data, not around preserving or governing human intent as a first-class input. That is not an oversight correctable with more funding or more urgency. It is what the system is for. A determination layer that governs what an Aptiv is authorized to do before it acts is a different kind of thing, built to a different purpose, than a model built to predict the next token more accurately.
Supporting a governed substrate is a specific, different ask from either halting AI or funding more oversight of it. It means directing attention, capital, and policy toward companies whose objective is not to build a smarter model, but to build the layer a model, of any capability level, has to pass through before its outputs become real-world actions. In this series' architecture, that is the difference between an Aptiv proposing intent and a governed layer determining whether that intent is authorized to execute. The model can keep getting more capable. What changes is that its capability no longer translates directly into unchecked action.
The determination layer in that diagram has to be an Era 3 substrate specifically, not a better-built Era 1 or Era 2 system. Era 1's procedural code has no way to take intent as a governing input at all; it executes instructions, and an instruction is not a determination about whether an action is authorized. Era 2's statistical systems, the frontier models this entire paper is about, produce outputs by inference, and inference is exactly the operation this series has spent two papers showing cannot also serve as its own constraint. A model asked to determine whether its own proposed action is authorized is still doing inference, just about a different question. Neither era has an architectural slot for determination as something separate from generation. Era 3, where intent is the computational primitive rather than a prompt fed into a statistical process, is the only one of the three where that separation is structural rather than aspirational. This is why "better guardrails" in Section 03 keeps producing the same failure regardless of how sophisticated the guardrail gets: a guardrail built on Era 1 or Era 2 architecture is still Era 1 or Era 2, no matter how it's marketed.
This reframes what "safety investment" should mean for each of the four constituencies from Section 01. For employees inside the labs, it means directing career capital toward governance-layer work specifically, since that is where the leverage described in Section 05 actually sits. For government officials, it means funding and mandating substrate-level standards, not just model-level audits. For investors, it means recognizing that a company solving determination is not a slower-growing, defensive bet against a faster-growing model company; it is infrastructure the faster company will eventually need in order to keep operating at all. For the public, it means directing pressure at "does this get built on a governed substrate," not "does this get built at all."
The natural objection to a substrate-first ask is that it sounds like the cautious, slower option, while the labs racing toward self-improving systems keep moving at full speed regardless. The validated performance data this series has cited elsewhere argues the opposite. Governing intent before execution has been measured, independently, at 20 to 114x acceleration and up to 99.7% energy reduction across tested workloads. A governed substrate is not a brake applied to a fast system. It is closer to a transmission: it is the thing that lets speed convert into useful, authorized motion instead of into failure modes nobody can catch in time.
This matters for the "stop versus don't stop" framing directly. A jurisdiction, a company, or an investor that funds substrate-layer governance is not choosing slower AI over faster AI. It is choosing AI whose speed is usable, because the layer determining what's authorized to happen doesn't have to wait for the model to finish inferring before it can act. That is a genuine third option, distinct from both halting the data centers and hoping the guardrails keep up.
Naming the right question only matters if it changes what each constituency actually does next.
The public should stop treating "pause AI" petitions and "ban AI" campaigns as the only available lever, and start asking which companies in the ecosystem they're supporting, as consumers and as citizens, are building governance rather than only capability. Employees inside frontier labs, the ones already saying privately what Coxon said publicly, have more leverage than a resignation post: internal advocacy for adopting or building a determination layer is a request their own leadership has now admitted, on the record, they don't currently have an answer to. Government officials drafting AI legislation, including the UK's Artificial Superintelligence Security Bill Hinton backed this week, should treat substrate-level governance standards as the enforceable target, not model-level behavior commitments that age out with every new release. Investors evaluating the sector should weight governance-layer platforms not as a hedge against frontier AI's success, but as a dependency of it, the piece of infrastructure the fastest-moving labs will need once "we don't have a plan" stops being an acceptable answer to a regulator, a court, or a customer.
It is not stop the data centers being built. It is not stop advancement in AI. It is support companies that are not trying to build smarter AI, but a better substrate for AI to operate on top of. It is not build better guardrails patched on as workarounds after the system is already fixed. It is support companies building the layer that determines, architecturally, what a system of any capability level is authorized to do. And it is not any substrate that happens to call itself governance. Section 06 was specific about this: it has to be an Era 3 substrate, because Era 1 and Era 2 architecture have no structural slot for determination separate from generation, no matter how the guardrail built on top of either one is marketed.
The public, employees, officials, and investors named in this paper are not wrong to be alarmed. Paper 67 showed the alarm is shared by the people building the technology itself, on a timeline this series had previously framed as generational and now has to correct: the people making this admission put the risk inside the next decade, not the next generation. They are asking the wrong two questions in response to that alarm, and demanding the wrong two actions. There is a third question, and it is the one this series has been answering since its first paper: not whether to stop the machine, but what it's standing on. That question is no longer one current generations get to leave to the next one to answer.