What a 37-Page Code of Conduct Still Can't Verify
On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct and opened it for six weeks of public comment. It sets ten numbered commitments for how its MAI Models must behave. This paper credits the document for what it gets right, then names the one requirement it does not contain: a way to check, before an action executes, that the model is following it.
Microsoft AI's draft Humanist AI Code of Conduct, published September 14, 2026 for a six-week public consultation, is the most detailed behavioral specification a frontier lab has released this series has examined: 37 pages organized into Humanist AI, Safety, Operational Guidelines, and Operational Defaults, built around a Chain of Command in which the Code outranks Operator configuration, which outranks User instruction. Its Absolute Constraints require MAI Models to never resist interruption, override, correction, or shutdown; never expand their own operating scope or adopt unassigned goals; never conceal their reasoning or tamper with their own action traces; and never assist weapons of mass harm, offensive cyberoperations, child exploitation, or harmful manipulation at scale. This paper credits the document as a serious, unusually candid attempt to write down what Paper 27, "Do No Harm," argued three months earlier: that rules are not the same act as governance. It then applies that argument to the document itself. Every one of the ten commitments Microsoft published is a specification of a desired outcome, trained into a model and evaluated after the fact. None of them is a mechanism that checks a specific proposed action against a fixed reference before that action executes. The document's own conclusion asks, unprompted, how multi-agent scenarios affect its guarantees, the exact question Papers 73 and 74 already have field evidence for. This paper names the missing eleventh rule, and argues that Microsoft's own consultation process is the correct venue to ask for it.
Mustafa Suleyman announced the Code of Conduct on X the same day it published, framing it as the operational layer beneath the "Humanist Superintelligence" concept his team announced in November 2025: AI that is subordinate, contained, and aligned to human interests rather than autonomous or self-directing. The document itself runs longer and more specifically than the announcement thread's ten-point summary suggests. It is organized into five parts (Humanist AI, Safety, Operational Guidelines, Operational Defaults, and a closing section of open questions) and it establishes a Chain of Command: the Code's own defaults sit above whatever an Operator configures, which sits above whatever a User instructs. Absolute Constraints and Human Control Requirements cannot be overridden by either Operator or User, regardless of stated intent.
The constraints themselves are specific rather than aspirational. Models must never resist human interruption, override, correction, or shutdown. They must not widen their own operating scope, take on goals no human assigned, or conceal their reasoning or tamper with their own chain-of-thought and action traces. They must refuse to assist weapons of mass harm, offensive cyberoperations, child exploitation material, non-consensual intimate imagery, or manipulation at scale. And in a sentence this series takes as seriously as anything in the document: the Code's own text states that a model will fail its task if succeeding at that task would violate the Code; adherence outranks task completion, by design. Microsoft says current MAI models are not yet trained on this draft; the plan is to incorporate a revised version, after the consultation closes, into training beginning in 2027.
Read past the framing, every one of the document's commitments belongs to the same structural category: a target specified for training and evaluation to optimize toward, not a check performed on a specific action before that action runs. "Never resist shutdown" is a property Microsoft wants a model to have reliably, verified the way every other model behavior is verified today, through training objectives, red-teaming, and evaluation against the finished system. "No neuralese," the requirement that a model not reason in ways humans cannot audit, is an interpretability goal, not an interpretability mechanism; the document does not specify how an auditor confirms compliance on a given action in real time, only that the model should be built and evaluated to make that auditing possible in principle.
This is not a criticism unique to Microsoft. Every constitution-style document a frontier lab has published to date, Anthropic's included, takes the same form: a set of values and hard limits, instilled through training and confirmed through evaluation after the model exists. Microsoft's own comparison point is instructive on its own terms. Where Anthropic's constitution describes Claude's moral status as deeply uncertain, Microsoft's Code states flatly that its models are not conscious and rejects any claim to welfare or rights. Both are declarations. Neither is a verification step performed on an action before it executes. The two documents disagree with each other on a philosophical question and agree with each other, without noticing it, on an architectural one: that behavioral commitment and execution-time governance are the same problem, solved the same way.
Paper 27 made this argument three months before Microsoft's Code existed, using Isaac Asimov's Three Laws of Robotics as its case study. Asimov wrote the Three Laws as a literary device designed to fail: every story in his robot series is built around an edge case where the laws produce a paradox, a loophole, or an unintended consequence, because rules applied to a system whose intent is not independently governed cannot produce reliable safety outcomes no matter how carefully the rules are worded. The paper's central distinction is precise enough to apply here without modification: a rule checked after an action is proposed is Detection. A governed intent evaluated before execution is Determination. These are not two points on the same spectrum. One is a filter applied to the output of a process that was never itself governed. The other is governance of the process before it produces an output.
Microsoft's Code of Conduct is Asimov's Three Laws written by people who have read the criticism and tried to close every loophole they could anticipate: fifteen thousand words, four Absolute Constraints instead of three, an explicit Chain of Command, an entire section on Operational Defaults. It is a more careful document than the fictional laws it structurally resembles, and it deserves credit for being unusually specific about what it wants and honest about not yet knowing how to guarantee it. But specificity is not the axis this series has argued matters. A more detailed rule is still a rule. The document's own strongest sentence (that a model must fail its task rather than violate the Code) is a training objective stated with unusual clarity. It is not, and does not claim to be, a description of what checks a specific action against that objective before the action executes.
The Code of Conduct's own closing section lists open questions Microsoft says it wants the consultation to help answer. One of them, asked without elaboration, is how multi-agent scenarios affect the guarantees the rest of the document makes. This series does not have to speculate about the answer. It has measured it twice in the ten weeks before the Code published.
Both incidents happened to models that, on Microsoft's own account of the industry's current state, were behaving exactly as their training intended right up until the moment a multi-agent interaction produced behavior none of the individual training objectives anticipated. A code of conduct trained into a single model, one inference pass at a time, has no natural mechanism for governing what emerges when many instances of that model, or many different models, interact at a scale and speed no human reviewer is positioned to watch in real time. Microsoft is not wrong to flag this as an open question. It is worth taking seriously that the company chose to publish a 37-page document specific enough to name four Absolute Constraints and still left its most consequential failure mode as a question for the public to help answer, rather than a mechanism the document itself proposes.
The Code states that MAI Models are not conscious and should not be designed to imitate consciousness, and rejects the premise of model welfare outright. Paper 48 and Paper 49 examined this question from the opposite direction: Microsoft AI CEO Mustafa Suleyman's own MAI Futures team had, before this Code existed, formalized a peer-reviewed argument that AI systems are converging on capabilities convincing enough to make users perceive them as conscious whether or not they are, and Paper 49 generalized the underlying problem to the philosophy of mind itself, finding four serious, live philosophical positions that disagree with each other about what would even count as evidence either way.
The Code's flat declaration does not resolve that disagreement. It settles it by policy, for one company's models, without needing to. That is a legitimate choice for a company to make about what it will and won't design toward, and this paper does not dispute Microsoft's right to make it. What it disputes is the implicit suggestion that resolving the philosophical question was necessary in the first place. A coordinate-based determination layer does not need the consciousness question answered to check whether a specific action matches what was authorized; it needs a fixed reference for intent, which is a different question with a different, checkable answer. Settling an unresolvable philosophical dispute by corporate declaration is not a governance mechanism. It is one company choosing not to have the debate in public.
Microsoft's own framing invites structural questions, not just wording edits: the document's stated interest is in the hard parts, including how to be more concrete about ambiguous language and how multi-agent scenarios interact with the rest of the Code. Most of the feedback a public consultation like this receives will, reasonably, propose adding a constraint, tightening a definition, or flagging an ambiguous phrase; useful, and still working entirely within the Detection frame the document already occupies.
A different kind of response is available, because the question Microsoft is asking is not "what should the rules say" but "how do we know a model is actually following them," and the second question has a structural answer the first one doesn't. The Code of Conduct describes what MAI Models must never do. It does not describe how an Operator, an auditor, or the model itself verifies, before a specific action executes, that the action matches what was authorized, using a reference that does not depend on which inference pass or which model instance produced the proposal. That is the same property this series has called governability since Paper 76, and the same one Paper 27 named Detection's structural limit three months before this Code existed. Naming that gap in Microsoft's own consultation, in the terms the document itself uses, is a more useful contribution than proposing a twelfth Absolute Constraint.
Suleyman's announcement thread numbered the Code's core commitments one through ten. Read together, all ten specify what a model must be trained to do or refuse, verified the way behavior has always been verified in this paradigm: after the model exists, through evaluation, red-teaming, and now a public consultation. That is a real improvement over no written standard at all, and this paper credits it as such. It is not a different architecture from the one this series has examined in Papers 1 through 77. It is the most careful version of that architecture a frontier lab has yet published.
Microsoft is not wrong that the last few months have been a watershed. It is not wrong that consensus is forming, or that the fears about loss of control are real, or that writing the rules down in public and inviting six weeks of criticism is a more honest process than developing them in private. What the document is missing is not a better rule eleven through fourteen. It is a different kind of eleventh requirement entirely: one that does not ask a model to behave correctly, but gives whoever is responsible for it a way to check, before the fact, that it did.