The Style Confusion Proof

Grounding Detection ≠ Determination in Measured Model Internals

An ICML 2026 paper set out to explain why prompt injection still works despite years of safety training. What it found is a mechanical account of a failure this series has argued from architecture alone for forty papers: models infer authority from how text sounds, not from any structural signal of where it came from, and the authors can show it, in the model's own activations, with numbers.

Ken Granville CEO & Co-Founder, MindAptiv White Paper 41 The Governed Machine August 2026
Abstract

Detection ≠ Determination has always rested on a distinction most AI discourse collapses without noticing: a model recognizing a pattern is not the same act as a system determining what should follow from it. This series has argued that distinction from architecture and fifteen years of platform work. What it lacked was a demonstration from inside a model's own activations that the collapse is real and measurable.

"Prompt Injection as Role Confusion" (Ye, Cui, and Hadfield-Menell, MIT, accepted ICML 2026) supplies it. The authors show that large language models infer the source of a piece of text (user, tool, system) from how it is written, not from the role tag wrapping it. Using internal measurement instruments they call role probes, they demonstrate that text imitating a trusted role's style comes to occupy the same representational space as text that actually holds that role.

Their proof-of-concept attack, CoT Forgery, injects a fabricated chain-of-thought trace styled to read as the model's own reasoning, achieving 60% attack success against frontier models on StrongREJECT and 61% in an agent-exfiltration setting. The paper's decisive result is a destyling test: rewriting the same forged content without its stylistic markers collapses attack success to roughly 10%, isolating style (not content, not tags, not argument quality) as the variable driving the model's perception of authority.

This paper grounds Paper 39's claim in those figures directly, and argues the finding is not a patchable gap in refusal training but a property of the same mechanism that lets these models write fluently at all, which is precisely why the fix has to be architectural, not behavioral.

Section 01The Claim, Restated

Detection ≠ Determination has always rested on a distinction most AI discourse collapses without noticing: a model recognizing a pattern is not the same act as a system determining what should follow from it. GenAI proposes. Synergy® governs. The doctrine has stood on architectural reasoning and fifteen years of platform work. What it lacked, until this paper, was a demonstration from inside a model's own activations that the collapse is real and measurable.

Ye, Cui, and Hadfield-Menell did not set out to confirm a governance doctrine. They set out to explain why prompt injection works, despite years of safety training aimed at preventing it. What they found is a mechanical account of exactly the failure this series has spent forty papers describing at the institutional and architectural level.

Section 02Role Confusion

The authors' central finding is that large language models do not identify the source of a piece of text by the role tag wrapping it: <system>, <user>, <tool>. They identify it by how the text sounds. A model treats every input it receives as a single undifferentiated stream, and infers “who is speaking” the way a reader infers an author's voice: from style, register, and rhythm, not from a label stapled to the outside.

To measure this directly rather than infer it from attack outcomes, the authors built what they call role probes: internal instruments trained only to read out how a model represents “who is speaking” at a given point in its processing, independent of whatever the model ultimately outputs. The finding is not that models are sometimes fooled by cleverly worded attacks. It is that injected text, once it convincingly imitates the style of a trusted role, occupies the same representational space inside the model as text that actually holds that role. There is no internal flag, no separate compartment, no structural distinction left to exploit. As far as the model's own perception is concerned, the imitation and the original are not close; they are the same thing.

Section 03CoT Forgery

The paper's proof-of-concept attack, CoT Forgery, injects a fabricated chain-of-thought trace into a user prompt or a tool output, styled to read exactly like the model's own internal reasoning. The model does not evaluate whether this reasoning is true. It recognizes the reasoning as its own, because it sounds like its own, and proceeds from there as if it had thought the thought itself.

The authors report average attack success of 60% against frontier models on the StrongREJECT benchmark, and 61% in an agent-exfiltration setting, against near-zero baselines when the same content is presented without the stylistic disguise. These are not edge-case jailbreaks requiring elaborate setup. They are a directly measurable consequence of how the model represents authorship internally, the mechanism Section 02 describes, made operational.

Reported Attack Success
60% (StrongREJECT, frontier models) · 61% (agent exfiltration) · near-zero baseline without stylistic disguise.

Section 04The Causal Proof

A skeptical reading of Sections 02 and 03 might grant that CoT Forgery works, while doubting that style specifically (as opposed to content, cleverness, or some other confound) is what's doing the work. The authors close that gap directly. When the same forged reasoning is “destyled” (rewritten to carry the identical content and logical structure, stripped of the stylistic markers that make it read like the model's own voice) attack success collapses from 61% to roughly 10%.

That collapse is the paper's load-bearing result. It isolates style, and only style, as the variable the model treats as a proxy for authority. Not the plausibility of the reasoning. Not the tag wrapping the text. Not even the content of the request itself: a further test the authors ran found that absurd justifications succeed at rates comparable to plausible ones, so long as both are styled correctly. The model is not evaluating whether an argument is good. It is checking whether an argument sounds like something it would have said, and treating a “yes” as sufficient grounds to act on it.

The Destyling Test
Same content. Same logical structure. Stylistic markers removed. Attack success: 61% → ~10%.

Section 05What This Confirms

None of this describes a bug that better prompting, more safety training, or a patch cycle will retire. Role confusion is not a gap in what these models have been taught to refuse. It is a property of how they represent authorship at all, the same representational layer that lets them write fluently in the first place. Teaching a model to be more suspicious of forged reasoning means teaching it to distrust its own internal voice, which is a much harder and much less stable ask than patching a known exploit string.

This is the empirical floor Paper 39 argued for and this paper now supplies with numbers: a system that infers authority from how something sounds cannot be the same system that determines whether that something should be permitted to happen. The two have to be architecturally separate, with the determination made by a layer that never mistakes a convincing style for a verified source.

What Ye, Cui, and Hadfield-Menell Confirmed
GenAI proposes. Synergy® governs.
Their activations proved the distinction. They did not need to write it.
The Governed Machine: Paper 41

A convincing style
was never a verified source.

Ye, Cui, and Hadfield-Menell set out to explain a security exploit and produced, as a side effect, a measurable proof of an architectural claim this series has made since Paper I. A model that grants authority based on how something sounds cannot also be the layer that determines whether it should be allowed to happen. That determination has to sit somewhere the style of the request was never the input.

Request Platform Access → Full White Paper Series

White Paper Series · The Governed Machine

1The Civilizational Fault Line 2We Are Building the Wrong Machine 3The Ornithopter Mistake 4The Convergence 5The Four Horsemen of the Knowledge Apocalypse 6What the Insiders Confirmed 7The Metaphor Trap 8The Recall Standard 9The $1 Trillion Governance Gap 10The Litigation Layer 11The Scale of Intent 12The Intent Economy 13The Session Illusion 14The Necessary Sequence 15The Wrong Race 16The Ledger That Is Intent-Driven 17The Agency Illusion 18The Substrate 19The End of the Mean 20Era 3: The Architecture of the Next Civilization 21The Missing Substrate 22The Context Fatigue Ceiling 23The Iceberg Stays Frozen 24The Dependency Tax 25The Record That Was Never Kept 26Composable by Default 27Do No Harm 28The Stack Replacement Thesis 29The Moat Is the Code 30The Last Platform War 31Beyond the Agent: Intent-Native Execution 32The Hardware Imagination 33The Architecture Tax 34The Tokenization Ceiling 35The Payment Moment 36The Oracle Problem 37The Reviewer Problem 38The Provenance Fallacy 39Role Without Determination 40Known and Funded Anyway 41The Style Confusion Proof ← this paper 42The Verification Tax 43The Pause Reflex 44The Human Margin 45The Balance of Power Fallacy 46The Liability Backstop 47One Substrate, Every Signal 48The Attribution Problem 49The Consciousness Ceiling 50The Detection Patch 51The Consumptive Machine 52The Agent That Isn't 53The Legibility Gap 54The Semiotic Machine 55The Transpilation Ceiling 56The Provisioning Ceiling 57The Reservation Ceiling 58The Circularity Ceiling 59The Coexistence Ceiling 60The Conformance Ceiling 61The Preservation Ceiling 62The Parity Clause 63The Governed Boundary 64The Transcript Problem 65The Unpaired System 66The Memory Ceiling 67The Admission Gap 68The Wrong Ask 69The Best Case 70The Last Chokepoint 71The Fourth Step 72The Adoption Standard 73The Same Weekend 74Sixty to One 75Coordinates, Not Correlations 76The Governability Axis 77Era 3, Confirmed 78The Eleventh Rule 79The Seventh Admission 80The Authorization Gap 81The Authorship Fallacy 82The Camera and the Vault 83Cleared to Proceed 84A Class, Not a Product 85The Inherited Playbook