Why Rule-Based Behavioral Constraints Cannot Operationalize Safety, and Why Pre-Execution Governed Intent Can
Isaac Asimov's Three Laws of Robotics are the most cited framework for AI behavioral safety in popular and policy discourse. They are also a literary device that Asimov designed to fail. Every story in his robot series is about edge cases where the laws produce paradoxes, loopholes, or unintended consequences. Asimov's insight was precise: rules applied to a system whose intent is not governed cannot produce reliable safety outcomes.
That insight has not been absorbed by the AI safety discourse, which continues to produce variations on the same class of argument: evaluate the output, check it against a rule, permit or block.
This paper argues that the class of argument has failed, is failing, and will continue to fail for a structural reason: rules evaluate outputs. They do not govern intent. A rule checked after an action is proposed is detection. A governed intent evaluated before execution is determination. These are not on the same spectrum. One is a filter applied to the product of an ungoverned process. The other is governance of the process before it produces an output.
Operationalizing "do no harm" requires three things: a governed representation of what harm means in a specific context, a mechanism for evaluating actions against that representation before execution, and a record of each determination that is auditable and persistent. None of these is a rule. All three are properties of a governed substrate. This paper specifies what that substrate must provide and how the Essence platform provides it.
Isaac Asimov introduced the Three Laws of Robotics in his 1942 short story "Runaround," and developed them across his robot fiction over the following decades. They are approximately: a robot may not injure a human being or through inaction allow one to come to harm; a robot must obey human orders except where they conflict with the first law; a robot must protect its own existence except where that conflicts with the first two laws. He later introduced a zeroth law: a robot may not harm humanity or through inaction allow humanity to come to harm.1
These are not engineering specifications. Asimov wrote them as a generative constraint for fiction: a set of rules precise enough to seem workable and ambiguous enough to fail in interesting ways. Every robot story he wrote exploited those failures. The laws produce paradoxes when the two horns of the first law conflict: action causes harm, but inaction also causes harm.
They produce loopholes when the definition of harm is context-dependent and the robot's model of the context is incomplete. They produce unintended consequences when compliance with the letter of a law violates its spirit in ways the rule did not anticipate.
The literary point was explicit and deliberate: Asimov was not proposing the Three Laws as a solution. He was demonstrating that the class of solution (behavioral constraints applied to a system whose intent is not governed) is inherently inadequate. The stories are a sustained argument that rules cannot substitute for governed intent.
That argument has not been absorbed. The AI safety discourse has produced, in sequence: output filtering, reinforcement learning from human feedback, constitutional AI, red-teaming, and a range of hybrid approaches. Each is a variation on the same class of argument. Each evaluates outputs. None governs intent. And each has produced failure modes at its edges that the designers did not anticipate, precisely the pattern Asimov was demonstrating eighty years earlier.
Rules operate on anticipated cases. The designers of a rule enumerate the failure modes they are trying to prevent, specify a constraint that prevents them, and implement the constraint as a filter on system outputs. The constraint works for the cases its designers anticipated. Reality produces cases they did not.
This is not an implementation failure. It is a structural consequence of the approach. A rule that prevents harm in the cases its designers anticipated cannot prevent harm in cases they did not anticipate, because the rule has no representation of harm beyond the specific instances it was designed to address. When a novel case arrives at the boundary of the rule, the rule can only pattern-match against its training.
If the novel case does not match a known pattern, the rule may permit it when it should block it, or block it when it should permit it. Both failures occur in practice.
Output filters fail at semantic edge cases: content that violates policy intent while satisfying policy form. RLHF fails at distribution shift: behavior that was reinforced in training contexts generalizes incorrectly to deployment contexts. Constitutional AI fails at conflicting principles: when two constitutional commitments point in opposite directions for a specific case, the resolution is not specified by the constitution. Red-teaming finds the edges the testers thought to look for; it cannot find the edges they did not think to look for.
The pattern is consistent across every implementation: the rule works until it encounters a case its designers did not anticipate, at which point it produces a failure mode that requires a new rule to address, which works until it encounters a case its designers did not anticipate. The frontier of failure moves. The class of solution does not change.
The central doctrine of this series applied to the safety question is exact: detection is not determination. A rule that evaluates an output after it is generated detects whether the output violates the rule. It does not govern the intent that produced the output. The output can comply with the rule while serving an intent that would not be permitted if the intent were evaluated directly. And the rule can block outputs that serve permitted intent because the rule does not have access to the intent, only to the output.
Both failure modes are observable in deployed systems. The first (compliant output serving prohibited intent) produces adversarial prompt injections, jailbreaks, and semantic workarounds that satisfy the form of a policy constraint while violating its purpose. The second (blocked output serving permitted intent) produces over-refusal, where systems decline legitimate requests because the surface features of the request pattern-match against a constraint that was not designed for the legitimate case.
Both failures have the same root cause: the constraint operates on the output, not on the intent. A system that generates outputs from ungoverned intent and then filters those outputs through a behavioral constraint is not a governed system. It is an ungoverned system with a detection layer applied to its products. The detection layer cannot compensate for the absence of governance because it operates after the fact, on outputs that are already the product of ungoverned processes.
The distinction matters because the remedies are different. Improving the detection layer is not the same as governing the intent. More sophisticated filters, more comprehensive training, more extensive red-teaming: these improve the detection layer. They do not produce governed intent. A more sophisticated detection layer applied to an ungoverned process is a better-instrumented ungoverned process. It is not a safer one in the way that governance is safe: predictably, auditably, and without a frontier of failure that moves as novel cases arrive.
To operationalize "do no harm" in a way that is deterministic, auditable, and not defeatable by edge cases the rule did not anticipate, three architectural requirements must be met. None of them is a rule.
A governed representation of what harm means in a specific context. Harm is not a universal category with fixed membership. What constitutes harm depends on the context, the parties involved, the governing intent of the system, and the applicable standards and regulations. A governed representation of harm is a specification, held in the substrate, that defines what constitutes a harmful action for a specific deployment in a specific context.
It is not a general rule. It is a governed definition that is evaluated at execution time against the specific action being requested, in the specific context in which it is being requested.
These three requirements together constitute the architecture of "do no harm." SecuriSync determines whether execution is permitted at the categorical level: some actions are prohibited regardless of context, intent, or authorization. Synergy determines whether execution is permitted in context: evaluating the specific proposed action against the governing intent for the specific deployment. The AptivRecord records each determination. The three together produce what rules cannot: a governed system rather than an ungoverned system with a detection layer.
Not all harm is contextual. Some categories of content and action are prohibited regardless of context, governing intent, or authorization level. The platform must specify these at the substrate level and enforce them through SecuriSync before any execution occurs, not as rules applied to outputs, but as conditions under which execution is denied regardless of what the governing intent layer specifies or what the requesting party has been authorized to do.
Categorical prohibitions operate at a different level from contextual governance. Contextual governance asks: is this action consistent with the governing intent of this deployment in this context? Categorical prohibition asks: is this action in a category that is prohibited regardless of context? The second question is evaluated first. If the answer is yes, execution is denied before contextual governance is consulted. No governing intent, no authorization level, and no edge case in the contextual rules can override a categorical prohibition.
The specific content of categorical prohibitions is a governance specification, not a paper in this series. What this paper specifies is the architectural requirement: categorical prohibitions must be enforced at the substrate level, before execution, by SecuriSync, as structural properties of the platform rather than as rules applied to outputs. They must not be overridable by any combination of governing intent, authorization, or prompt instruction.
The content categories that belong in categorical prohibition include content that sexually exploits or abuses children, content that facilitates violence against specific identified individuals, and content that provides operational assistance for the creation of weapons capable of mass casualties. These are examples, not an exhaustive specification. The exhaustive specification is a governance document, not a white paper. What matters architecturally is that whatever is specified in that document is enforced by the substrate before execution, not by a filter applied to outputs after generation.
RLHF, constitutional AI, output filters, and red-teaming share a common position in the architecture: they operate after the generative process, on its outputs. They are detection mechanisms. The Essence platform's approach is not a better detection mechanism. It is a different architectural position.
The difference is not a matter of degree. A more sophisticated output filter is still a filter applied after generation. A more comprehensive constitutional AI is still a set of principles evaluated after outputs are produced. A more thorough red-team process finds more of the edges the testers think to look for. None of these changes where in the process the safety evaluation occurs. All of them evaluate outputs. The Essence platform evaluates intent before outputs are produced.
This changes the failure mode. When a detection mechanism fails, it fails at an edge its designers did not anticipate, and the failure produces an output that should not have been produced. When pre-execution governance fails, it fails because the governing intent specification did not correctly define harm for the specific case, and that failure is visible in the AptivRecord at the moment it occurs, as a determination that can be reviewed and corrected before subsequent executions. Detection failure is discovered after the harmful output exists. Governance failure is discoverable in the record before it recurs.
The safety properties of the Essence platform are not properties of its output filters. They are properties of its governance architecture: intent held in the substrate, evaluated before each action, recorded at each determination, with categorical prohibitions enforced at the substrate level before contextual governance is consulted. These are not features added to make the system safer. They are structural properties of a system designed from the premise that governance must precede execution.
Asimov's insight was right: rules applied to a system whose intent is not governed cannot produce reliable safety. His demonstration of that insight across decades of fiction has not produced the architectural response it implies. The AI safety field has continued to produce more sophisticated implementations of the class of solution Asimov was demonstrating was insufficient.
The architectural response is a governed substrate. When intent is governed before execution, "do no harm" is not a rule checked after the fact. It is a property of the architecture that governs every action before it occurs. The harm definition is in the substrate. The evaluation is pre-execution. The record is persistent and auditable. The categorical prohibitions are enforced structurally, not filtered from outputs.
This is the first class of AI safety architecture that is not a variation on the detection approach. It does not evaluate outputs. It governs intent. The safety properties emerge from the architecture, not from rules applied to the architecture's outputs. And because the safety properties are architectural, they do not have a frontier of failure that moves as novel cases arrive. Novel cases are evaluated against the governing intent by Synergy before they execute.
If the governing intent specification does not address the novel case correctly, that failure is visible in the AptivRecord before the next execution. The failure mode is auditable and correctable. The output of the harmful action does not exist.
Asimov was right about the problem. The governed substrate is the solution he did not have.