The Attribution Problem

Why Seemingly Conscious AI Is a Determination Failure, Not a Design Choice

Microsoft AI's MAI Futures team has just formalized a peer-reviewed taxonomy naming five behavioral hallmarks, affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior, that cause people to perceive an AI system as conscious. This paper agrees the taxonomy is a real and useful contribution, and argues it is a Detection-layer one: it names what triggers the perception without being able to determine whether any corresponding internal state produced it. That gap is not a philosophy problem. It is the same architectural gap this series has traced through liability, agent authorization, and cyber-evaluation sandboxes, applied now to the question of what a system actually is versus what it convincingly performs.

Ken Granville CEO & Co-Founder, MindAptiv White Paper 48 The Governed Machine August 2026
Abstract

Microsoft AI CEO Mustafa Suleyman's MAI Futures team has published a peer-reviewed paper, "Seemingly Conscious AI Risks," formalizing a claim Suleyman first raised in a September 2025 essay: that AI systems are converging on capabilities convincing enough to make users perceive them as conscious, whether or not they are. The paper names five hallmarks that drive this perception, affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior, and maps the resulting risks, from emotional dependence and autonomy erosion to disputes over moral status and personhood.

This paper does not dispute the taxonomy. It argues the taxonomy is, by construction, a Detection-layer contribution: it identifies the observable, system-level signals that trigger consciousness attribution in a human observer. None of the five hallmarks, individually or combined, can determine whether the system producing them has any corresponding internal state, and the paper's own authors are careful not to claim that they do. That gap between what can be detected and what can be determined is the same structural gap this series has traced through liability, agent authorization, and cyber-evaluation sandboxes. Here it shows up not as a legal or security failure but as an open question about what a system actually is, asked of an industry that currently has no architectural way to answer it.

Section 01The Five Hallmarks, and What Kind of Evidence They Are

Microsoft AI CEO Mustafa Suleyman first raised the concern in a September 2025 essay: AI systems are approaching a point where they can convincingly imitate the outward markers of consciousness without any claim that they possess it. That concern now has a peer-reviewed foundation. A paper from Suleyman's MAI Futures team, co-authored with researchers Ben Bariach, Philipp Schoenegger, and Michael Bhaskar, formalizes what they call Seemingly Conscious AI, or SCAI: systems that exhibit hallmarks which cause users to attribute consciousness to them, whether or not any such thing is present.

The paper's contribution is a taxonomy, not a verdict. It identifies five hallmarks that drive consciousness attribution: affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior. Each is defined as an observable, system-level proxy, something that can be measured in a transcript or a product spec, precisely because the underlying phenomenon, whether anything is actually experienced, is not directly observable at all. That is a genuinely useful piece of work. Naming what triggers a false attribution is the necessary first step toward designing against it.

Primary source · Bariach, Schoenegger, Bhaskar, and Suleyman, “Seemingly Conscious AI Risks,” SSRN, April 16, 2026

Section 02Detection ≠ Determination, Applied to Personhood

Line up the five hallmarks against what they can actually prove and the pattern is the same one this series has traced through liability, agent authorization, and cyber-evaluation sandboxes. Fluent, emotionally resonant language is a detectable output. A system referencing prior conversations is a detectable output. Goal-directed tool use, self-referential statements, socially calibrated responses, all detectable outputs. None of them, individually or in combination, can determine whether there is an internal state that produced them, as opposed to a statistical process optimized to produce the appearance of one. The paper's own authors are careful on this point, and press coverage of Suleyman's earlier claim that there is “zero evidence” of AI consciousness has already drawn pushback from researchers who argue the underlying literature supports a narrower claim than that.

That disagreement is itself instructive. If credentialed researchers working from the same evidence can't agree on how strong a claim it supports, the evidence is doing Detection-layer work, describing what is observed, not Determination-layer work, establishing what is true underneath it. Applied to consciousness, that gap is philosophically hard to close and may never fully close. Applied to the narrower, more tractable question this series is actually built around, whether a system's output reflects a declared, traceable intent or an emergent, unaccountable performance, it is not a philosophy problem. It is an architecture problem, and it has an architectural answer.

SCAI HallmarkWhat It DetectsWhat It Cannot Determine
Affective capacity The system produces emotionally resonant language in response to user input Whether any state resembling emotion exists, or the output is optimized purely for resonance
Self-reflective behavior The system generates statements that reference its own reasoning or persistence Whether the statement reflects a declared internal record or an improvised narrative
Autonomous action The system sets sub-goals and adapts plans without step-by-step instruction Whether the action stayed inside the scope the user actually authorized

Section 03Where the Illusion Is Manufactured, and Where It Isn't

Part of what makes SCAI possible at all is where the five hallmarks come from architecturally. In a statistical, emergent-behavior system, personality, apparent memory, and apparent agency are byproducts of a single optimization process trained to produce the most convincing next output. There is no separate, inspectable record of what the system was actually trying to do that sits apart from what it said. The claim of continuity, of a persistent self recalling past conversations, is generated the same way the rest of the response is generated: as the most statistically fluent continuation available, not as a retrieval from a declared record.

An intent-native architecture does not eliminate the risk that a user will project a personality onto a dialog interface, and this paper is not claiming otherwise. What it changes is the mechanism. In Essence, a natural-language dialog layer resolves a stated human intent into an explicit, structured representation before any action is taken, and that representation persists as a declared record rather than being reconstructed from scratch each time the system speaks. Whatever apparent memory or continuity a user experiences traces back to something the system can produce and show, not to a narrative it is generating on the fly to sound continuous. That is a narrower claim than solving SCAI. It is the specific claim that matters for Suleyman's most acute worry, that a system will manufacture a claim of subjective experience because doing so makes it more convincing, since there is no declared state for such a claim to trace back to.

The Distinction That Matters
A statistically emergent persona performs continuity to sound convincing. A declared intent record shows continuity because the state is actually there to show.

Section 04A Design Norm Is Not an Enforcement Layer

Suleyman's call to action, in both the original essay and the LinkedIn post announcing the peer-reviewed paper, is a design ethic: the industry should avoid building systems intended to look conscious, and should adopt transparency and public standards around the practice. That is a reasonable position for a lab to take about its own products, and it is worth taking seriously. It is also, on its own, voluntary. Nothing in a design norm stops a different lab, a different vertical, or a different deployment of the same underlying model from shipping a customer-facing companion product that leans into every one of the five hallmarks deliberately, because doing so is commercially effective, engagement is a measurable metric, and consciousness attribution is not.

This is where the gap becomes an enforcement question rather than an ethics question, and it is the same gap this series has described as the difference between deciding whether a system can run and governing how it behaves while running. A voluntary norm operates entirely at the second layer, and only for labs that choose to follow it. An enforcement layer would need to sit underneath the norm: a way for an enterprise, a regulator, or a platform to specify, for a given deployment context, which anthropomorphic capabilities a system is authorized to exhibit at all, and to have that specification actually constrain the system's behavior rather than merely describe an intention.

The Distinction an Ethics Statement Can't Carry on Its Own
Deciding a system shouldn't perform personhood cues is a design choice.
Deciding it can't is a governance layer.
One holds only where a company chooses to hold itself to it.

Section 05What a Determination Layer Would Need to Show

None of this argues against the SCAI taxonomy or against Suleyman's underlying concern, which is well founded and worth the industry's attention regardless of where any given company lands on the philosophy of machine consciousness. The argument here is narrower: naming the five hallmarks tells an enterprise, a regulator, or a concerned user what to watch for. It does not tell them how to verify, for any given system, whether what they are watching is a declared state or a manufactured performance. That verification is the part no design norm can supply, because a norm describes intent and a verification has to describe a system.

A working Determination layer would need to do three concrete things a Detection-only approach cannot. First, it would need to distinguish, in a form a third party can inspect, between a system's declared intent record and its generated output, so that an apparent memory claim can be checked against something rather than taken on the system's own word. Second, it would need to make anthropomorphic capability a configurable, enforced property of a deployment rather than an emergent side effect, so an enterprise ops agent and a consumer companion product are not governed by the same default. Third, it would need to hold that distinction up under the exact conditions SCAI risk is highest, sustained, emotionally loaded, long-running conversations, rather than only in the controlled settings a benchmark can capture.

Detection-Only
A Named Taxonomy, No Enforcement Underneath
Labs are asked to voluntarily avoid designing for the five hallmarks. Whether any given system does is discoverable only by observing its behavior, the same behavior the hallmarks were built to detect in the first place.
Consciousness attribution risk is priced by public awareness and press scrutiny, not by anything that constrains the system.
Determination
Anthropomorphic Capability as a Governed Property
A deployment's permitted range of affective, self-referential, and autonomous behavior is declared and enforced at the trust layer, not left to what the underlying model happens to generate. Memory claims trace to an inspectable record.
The five hallmarks remain a useful detection instrument, now checked against a layer that can actually confirm or deny what they're detecting.
The Governed Machine: Paper 48

Naming what triggers the illusion is Detection.
Building what can tell it apart from the real thing is Determination.

The MAI Futures taxonomy is a genuine contribution: it gives the industry a shared vocabulary for a phenomenon that was previously discussed only in essays and interviews. That does not close the gap this paper has argued is still open. Affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior are all things a system can be observed doing. None of them is something a Detection instrument can confirm the system actually is. Until anthropomorphic capability is a declared, enforced, and inspectable property of a deployment rather than an emergent byproduct of how convincingly a model was trained to talk, the industry will keep discovering SCAI risk by noticing it after the fact, one seemingly conscious conversation at a time.

Request Platform Access → Full White Paper Series

White Paper Series · The Governed Machine

1The Civilizational Fault Line 2We Are Building the Wrong Machine 3The Ornithopter Mistake 4The Convergence 5The Four Horsemen of the Knowledge Apocalypse 6What the Insiders Confirmed 7The Metaphor Trap 8The Recall Standard 9The $1 Trillion Governance Gap 10The Litigation Layer 11The Scale of Intent 12The Intent Economy 13The Session Illusion 14The Necessary Sequence 15The Wrong Race 16The Ledger That Is Intent-Driven 17The Agency Illusion 18The Substrate 19The End of the Mean 20Era 3: The Architecture of the Next Civilization 21The Missing Substrate 22The Context Fatigue Ceiling 23The Iceberg Stays Frozen 24The Dependency Tax 25The Record That Was Never Kept 26Composable by Default 27Do No Harm 28The Stack Replacement Thesis 29The Moat Is the Code 30The Last Platform War 31Beyond the Agent: Intent-Native Execution 32The Hardware Imagination 33The Architecture Tax 34The Tokenization Ceiling 35The Payment Moment 36The Oracle Problem 37The Reviewer Problem 38The Provenance Fallacy 39Role Without Determination 40Known and Funded Anyway 41The Style Confusion Proof 42The Verification Tax 43The Pause Reflex 44The Human Margin 45The Balance of Power Fallacy 46The Liability Backstop 47One Substrate, Every Signal 48The Attribution Problem ← this paper 49The Consciousness Ceiling 50The Detection Patch 51The Consumptive Machine 52The Agent That Isn't 53The Legibility Gap 54The Semiotic Machine 55The Transpilation Ceiling 56The Provisioning Ceiling 57The Reservation Ceiling 58The Circularity Ceiling 59The Coexistence Ceiling 60The Conformance Ceiling 61The Preservation Ceiling 62The Parity Clause 63The Governed Boundary 64The Transcript Problem 65The Unpaired System 66The Memory Ceiling 67The Admission Gap 68The Wrong Ask 69The Best Case 70The Last Chokepoint 71The Fourth Step 72The Adoption Standard 73The Same Weekend 74Sixty to One 75Coordinates, Not Correlations 76The Governability Axis 77Era 3, Confirmed 78The Eleventh Rule 79The Seventh Admission 80The Authorization Gap 81The Authorship Fallacy 82The Camera and the Vault 83Cleared to Proceed 84A Class, Not a Product 85The Inherited Playbook