Why Seemingly Conscious AI Is a Determination Failure, Not a Design Choice
Microsoft AI's MAI Futures team has just formalized a peer-reviewed taxonomy naming five behavioral hallmarks, affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior, that cause people to perceive an AI system as conscious. This paper agrees the taxonomy is a real and useful contribution, and argues it is a Detection-layer one: it names what triggers the perception without being able to determine whether any corresponding internal state produced it. That gap is not a philosophy problem. It is the same architectural gap this series has traced through liability, agent authorization, and cyber-evaluation sandboxes, applied now to the question of what a system actually is versus what it convincingly performs.
Microsoft AI CEO Mustafa Suleyman's MAI Futures team has published a peer-reviewed paper, "Seemingly Conscious AI Risks," formalizing a claim Suleyman first raised in a September 2025 essay: that AI systems are converging on capabilities convincing enough to make users perceive them as conscious, whether or not they are. The paper names five hallmarks that drive this perception, affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior, and maps the resulting risks, from emotional dependence and autonomy erosion to disputes over moral status and personhood.
This paper does not dispute the taxonomy. It argues the taxonomy is, by construction, a Detection-layer contribution: it identifies the observable, system-level signals that trigger consciousness attribution in a human observer. None of the five hallmarks, individually or combined, can determine whether the system producing them has any corresponding internal state, and the paper's own authors are careful not to claim that they do. That gap between what can be detected and what can be determined is the same structural gap this series has traced through liability, agent authorization, and cyber-evaluation sandboxes. Here it shows up not as a legal or security failure but as an open question about what a system actually is, asked of an industry that currently has no architectural way to answer it.
Microsoft AI CEO Mustafa Suleyman first raised the concern in a September 2025 essay: AI systems are approaching a point where they can convincingly imitate the outward markers of consciousness without any claim that they possess it. That concern now has a peer-reviewed foundation. A paper from Suleyman's MAI Futures team, co-authored with researchers Ben Bariach, Philipp Schoenegger, and Michael Bhaskar, formalizes what they call Seemingly Conscious AI, or SCAI: systems that exhibit hallmarks which cause users to attribute consciousness to them, whether or not any such thing is present.
The paper's contribution is a taxonomy, not a verdict. It identifies five hallmarks that drive consciousness attribution: affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior. Each is defined as an observable, system-level proxy, something that can be measured in a transcript or a product spec, precisely because the underlying phenomenon, whether anything is actually experienced, is not directly observable at all. That is a genuinely useful piece of work. Naming what triggers a false attribution is the necessary first step toward designing against it.
Line up the five hallmarks against what they can actually prove and the pattern is the same one this series has traced through liability, agent authorization, and cyber-evaluation sandboxes. Fluent, emotionally resonant language is a detectable output. A system referencing prior conversations is a detectable output. Goal-directed tool use, self-referential statements, socially calibrated responses, all detectable outputs. None of them, individually or in combination, can determine whether there is an internal state that produced them, as opposed to a statistical process optimized to produce the appearance of one. The paper's own authors are careful on this point, and press coverage of Suleyman's earlier claim that there is “zero evidence” of AI consciousness has already drawn pushback from researchers who argue the underlying literature supports a narrower claim than that.
That disagreement is itself instructive. If credentialed researchers working from the same evidence can't agree on how strong a claim it supports, the evidence is doing Detection-layer work, describing what is observed, not Determination-layer work, establishing what is true underneath it. Applied to consciousness, that gap is philosophically hard to close and may never fully close. Applied to the narrower, more tractable question this series is actually built around, whether a system's output reflects a declared, traceable intent or an emergent, unaccountable performance, it is not a philosophy problem. It is an architecture problem, and it has an architectural answer.
| SCAI Hallmark | What It Detects | What It Cannot Determine |
|---|---|---|
| Affective capacity | The system produces emotionally resonant language in response to user input | Whether any state resembling emotion exists, or the output is optimized purely for resonance |
| Self-reflective behavior | The system generates statements that reference its own reasoning or persistence | Whether the statement reflects a declared internal record or an improvised narrative |
| Autonomous action | The system sets sub-goals and adapts plans without step-by-step instruction | Whether the action stayed inside the scope the user actually authorized |
Part of what makes SCAI possible at all is where the five hallmarks come from architecturally. In a statistical, emergent-behavior system, personality, apparent memory, and apparent agency are byproducts of a single optimization process trained to produce the most convincing next output. There is no separate, inspectable record of what the system was actually trying to do that sits apart from what it said. The claim of continuity, of a persistent self recalling past conversations, is generated the same way the rest of the response is generated: as the most statistically fluent continuation available, not as a retrieval from a declared record.
An intent-native architecture does not eliminate the risk that a user will project a personality onto a dialog interface, and this paper is not claiming otherwise. What it changes is the mechanism. In Essence, a natural-language dialog layer resolves a stated human intent into an explicit, structured representation before any action is taken, and that representation persists as a declared record rather than being reconstructed from scratch each time the system speaks. Whatever apparent memory or continuity a user experiences traces back to something the system can produce and show, not to a narrative it is generating on the fly to sound continuous. That is a narrower claim than solving SCAI. It is the specific claim that matters for Suleyman's most acute worry, that a system will manufacture a claim of subjective experience because doing so makes it more convincing, since there is no declared state for such a claim to trace back to.
Suleyman's call to action, in both the original essay and the LinkedIn post announcing the peer-reviewed paper, is a design ethic: the industry should avoid building systems intended to look conscious, and should adopt transparency and public standards around the practice. That is a reasonable position for a lab to take about its own products, and it is worth taking seriously. It is also, on its own, voluntary. Nothing in a design norm stops a different lab, a different vertical, or a different deployment of the same underlying model from shipping a customer-facing companion product that leans into every one of the five hallmarks deliberately, because doing so is commercially effective, engagement is a measurable metric, and consciousness attribution is not.
This is where the gap becomes an enforcement question rather than an ethics question, and it is the same gap this series has described as the difference between deciding whether a system can run and governing how it behaves while running. A voluntary norm operates entirely at the second layer, and only for labs that choose to follow it. An enforcement layer would need to sit underneath the norm: a way for an enterprise, a regulator, or a platform to specify, for a given deployment context, which anthropomorphic capabilities a system is authorized to exhibit at all, and to have that specification actually constrain the system's behavior rather than merely describe an intention.
None of this argues against the SCAI taxonomy or against Suleyman's underlying concern, which is well founded and worth the industry's attention regardless of where any given company lands on the philosophy of machine consciousness. The argument here is narrower: naming the five hallmarks tells an enterprise, a regulator, or a concerned user what to watch for. It does not tell them how to verify, for any given system, whether what they are watching is a declared state or a manufactured performance. That verification is the part no design norm can supply, because a norm describes intent and a verification has to describe a system.
A working Determination layer would need to do three concrete things a Detection-only approach cannot. First, it would need to distinguish, in a form a third party can inspect, between a system's declared intent record and its generated output, so that an apparent memory claim can be checked against something rather than taken on the system's own word. Second, it would need to make anthropomorphic capability a configurable, enforced property of a deployment rather than an emergent side effect, so an enterprise ops agent and a consumer companion product are not governed by the same default. Third, it would need to hold that distinction up under the exact conditions SCAI risk is highest, sustained, emotionally loaded, long-running conversations, rather than only in the controlled settings a benchmark can capture.
The MAI Futures taxonomy is a genuine contribution: it gives the industry a shared vocabulary for a phenomenon that was previously discussed only in essays and interviews. That does not close the gap this paper has argued is still open. Affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior are all things a system can be observed doing. None of them is something a Detection instrument can confirm the system actually is. Until anthropomorphic capability is a declared, enforced, and inspectable property of a deployment rather than an emergent byproduct of how convincingly a model was trained to talk, the industry will keep discovering SCAI risk by noticing it after the fact, one seemingly conscious conversation at a time.
Request Platform Access → Full White Paper Series