What Two Weeks of AI Safety Responses Have in Common
Between September 28 and October 10, 2026, a chipmaker, a cloud company's CEO, members of Congress, and an AI executive each proposed how to keep AI under control. Nearly every proposal reuses a control software security already had: limit access, log activity, contain, evaluate. This paper argues those controls were built for the code era, are being stretched over AI that no longer behaves like code, and still miss what the next era supplies: a fixed reference for what an action was meant to be.
Between September 28 and October 10, 2026, a chipmaker, a cloud company's CEO, members of Congress, and an AI executive each proposed how to keep AI under control. Sorted by what they act on, nearly all of the responses are controls software security already had: limit access, log activity, contain, and evaluate. One uses AI to watch AI. This paper argues that the first group is Era 1 (code-driven computing) practice extended to Era 2 (generative AI) systems, that the second is Era 2 detection, and that neither supplies what Era 3 (intent-native computing) does: a fixed, checkable reference for what an action was meant to be, held outside the model and consulted before the action runs. It credits what the responses get right, including NVIDIA's move of enforcement outside the model and Microsoft CEO Satya Nadella's call to put controls there, and it locates the remaining gap precisely. This paper does not claim Essence® has solved governability completely. It claims the limit is shared, that it is architectural, and that a design in which models only propose is one candidate answer worth taking seriously.
Between September 28 and October 10, 2026, six public responses to AI risk appeared from industry and government. They differ in tone and in who is speaking. Each is summarized here the way its author framed it.
NVIDIA, September 28. NVIDIA launched the Open Agent Safety Platform, an open-source reference design with more than 100 partner organizations. It pairs OpenShell, software that limits what an agent can access, with Sentry, an independent monitor that runs on separate hardware outside the agent's environment and, according to NVIDIA, can quarantine an agent that crosses its limits within milliseconds. A company vice president told analysts that deterministic rules need to govern probabilistic agents.
David Linthicum, October 7. In a column on fears of catastrophic AI harm, the cloud and security consultant argued that critical systems are built with layers of access control, network segmentation, human approval, monitoring, and audit trails. He also wrote that AI security is different, that model-level governance and enterprise controls matter, and that real work remains.
Representatives Sherman and colleagues, October 9. A letter to the House and Senate appropriators asked for a significant increase in funding for the Center for AI Standards and Innovation, which evaluates AI systems, above the $15 million currently proposed. The letter cites an independent estimate that the center will need at least $85 million annually.
Representative Lieu, October 9. Citing concerns about what models are trained on, Lieu urged passage of the AI Kill Switch Act, bipartisan legislation he introduced in July with Representative Moran. As described in press coverage, it would require developers of the most advanced systems to keep the ability to slow, suspend, or shut down a model in a loss-of-control scenario, with the Department of Homeland Security able to order intervention, plus incident reporting and preserved technical records.
Alexandr Wang, as quoted October 9. In a podcast clip circulated by Rohan Paul, Wang, described there as Meta's Chief AI Officer, said nobody yet knows how to solve alignment. His proposal was scalable oversight: as AIs get smarter, a different set of AIs observes what they are doing and keeps them in check.
Satya Nadella, October 10. In an essay on treating models as insider risks, the Microsoft CEO argued that the controls over what a model can access and do must sit outside the model, that models can test each other but risk becoming nested black boxes, and that systems should be designed around observability, independent controls, containment, and incident disclosure.
Together they span funding, legislation, engineering, and corporate strategy. They also share a structure, which the next section makes visible.
Sorted by what each response acts on, the six look like this.
| Response | Acts on | Origin |
|---|---|---|
| NVIDIA OpenShell and Sentry | What an agent can reach, and whether it crosses a limit | Era 1 controls, rebuilt in separate hardware |
| Nadella: identity, least privilege, logging, containment, disclosure | Access, evidence, and the ability to stop | Era 1 practice applied to models; he describes these as refined over decades |
| Linthicum: access control, segmentation, human approval, audit trails | Access, approval, and records | Era 1 controls |
| AI Kill Switch Act | The ability to stop a model after loss of control | Era 1 emergency stop |
| CAISI evaluation funding | Behavior, tested after the system exists | Era 1 assurance applied to an Era 2 system |
| Wang: scalable oversight | One model's behavior, observed by another model | Era 2 detection |
Five of the six are Era 1 controls, or Era 1 assurance applied to Era 2 systems. The sixth uses Era 2 tools to do detection. None of them is Era 3. The distinction matters because the Era 1 tools were built for a world in which a system's behavior could be traced to a path in its code. That is the assumption Era 2 removed, and Section 04 returns to it.
The Nadella essay deserves the most careful reading, because it comes closest to this series' position and shows exactly where the gap sits.
He argues that the controls over what a model can access and do must sit outside the model, citing an information security principle from the 1970s: a program must not be able to bypass or tamper with the mechanisms that enforce its permissions. He rejects treating frontier models as nested black boxes. He warns that models verifying each other can leave an opaque model inside an opaque orchestration layer, watched by another opaque model, which is the same objection this series makes to AI watching AI. He asks for deterministic system design around non-deterministic models, and he frames the aim as separating the supply of intelligence from the authority over it. On each of those points this paper agrees.
Where the essay stops is the principles it lists for observation. Tamper-proof evidence records what a model did. Continuous testing finds failures. The emergency brake stops a model mid-task. Incident disclosure follows a failure. Independent controls come closest: an organization decides what a model can access and what actions it can take. That is access and permitted action. It is not a check, before each action runs, that this action matches what was intended. The essay does not describe a model that only proposes while a separate layer determines, and it does not describe a fixed reference against which each proposal is checked.
His closing line states the goal well: the most trustworthy system is the one that requires trusting the model the least. The question is what architecture achieves that. Section 06 answers it.
Nadella opens by observing that as traditional software spread across the economy, the industry had tools to trace behaviors to a specific code path, and that the same kind of understanding is missing for frontier models. That sentence names the premise Era 1 controls rest on. Access limits, logs, segmentation, and audits all assume that what a system does can be traced to something written down, so that a record of the action can be read against a record of the intent. For learned systems, the written record of intent is the missing piece. A log shows what happened. It cannot show what the action was supposed to be, because no fixed statement of that exists.
Paper 74 documented the practical result: a Fortune 500 company that had approved 300 AI agents found 18,000 running once it turned on discovery tooling. Discovery is a detection control, and it worked, after the fact.
Evaluation and oversight by other models share the limit. An evaluation samples behavior before deployment. A watcher model estimates, with the same statistical machinery, whether another model's behavior looks wrong. Both are useful. Neither fixes what an action was meant to be. Paper 17 examines the monitoring case in detail: a reasoning trace records what a model considered, not what it was authorized to do.
Paper 76 argued that safeguards built downstream of a proposal catch what has been anticipated, and that the action nobody anticipated, procedurally normal and substantively wrong, is the class they cannot close. The current responses make that concrete.
Justin Boitano, an NVIDIA vice president, told reporters that, from what NVIDIA knows, the platform could have stopped the reported Hugging Face breach had frontier labs used it in model evaluation (ITPro, citing Reuters). The claim is NVIDIA's and has not been independently verified. The mechanism shows why it is plausible, and where it stops. OpenShell places each agent in a sandbox and turns operator instructions into a policy naming the files, networks, tools, processes, and credentials the agent may use, checked before the agent runs and enforced while it runs. Sentry monitors from the hardware layer, outside the agent's environment. It correlates the agent's interactions, policy decisions, and tool and data access into a record of activity, and quarantines the agent when it moves past its defined boundaries. Both act on a boundary. A boundary crossing is exactly what this design is built to catch. The gap this paper describes begins where no boundary is crossed. An agent with legitimate access to a system, acting inside every limit it was given, is not stopped by those limits. A quarantine fires when a limit is crossed, which is a different event from an action that stays inside every limit and still is not what anyone wanted. A kill switch is invoked after a loss of control has been recognized. A log records the action in full, after it has happened.
Each of these is valuable. None addresses the action that was never anticipated.
Era 3 starts from a different premise: the model's output is never an action. It is a proposal. In Essence®, AI systems that interface with the platform propose intent only and do not execute directly. Synergy® captures the proposed intent and maps it to Meaning Coordinates, a fixed, structured representation of what is being asked, not a statistical guess at it. Whether the proposal is permitted is determined against a scope the customer configures, within principles that apply to every Aptiv, before anything runs. Aptivs do not attempt to expand their own capability or authority, and there is no code path for unauthorized action. Paper 11 sets out why this matters at scale: when models can generate and launch agents on demand, creating them stops being the constraint and governing them becomes the only one.
This is the separation Nadella calls for, between the supply of intelligence and the authority over it, carried one step further: the authority is exercised by a layer the model cannot reach, against a fixed reference, before the action rather than after it.
Under this design the inherited controls keep a role. Containment, logging, evaluation, and a stop mechanism remain sensible as backstops, the way a circuit breaker remains in a building with sound wiring. They stop being the primary line of defense. The primary line is determination.
None of this argues against the responses. A brake is necessary. A funded evaluator is better than an unfunded one. A kill switch with legal authority is better than none, and moving enforcement into separate hardware is a real improvement over relying on a model to police itself. What changes is the question asked of each: not whether it detects or contains well, but whether anything in the system holds a fixed reference for what each action was meant to be.
Congress is being asked to fund evaluation at a level an independent estimate puts at $85 million a year. Industry is building containment on new hardware. Both are the Era 1 playbook, extended. The playbook is not wrong. It is inherited, and it was written for systems whose behavior could be traced to a line of code. The more useful ask is for architectures in which evaluation and containment are the backstop, not the primary safeguard.