Grounding Detection ≠ Determination in Measured Model Internals
An ICML 2026 paper set out to explain why prompt injection still works despite years of safety training. What it found is a mechanical account of a failure this series has argued from architecture alone for forty papers: models infer authority from how text sounds, not from any structural signal of where it came from, and the authors can show it, in the model's own activations, with numbers.
Detection ≠ Determination has always rested on a distinction most AI discourse collapses without noticing: a model recognizing a pattern is not the same act as a system determining what should follow from it. This series has argued that distinction from architecture and fifteen years of platform work. What it lacked was a demonstration from inside a model's own activations that the collapse is real and measurable.
"Prompt Injection as Role Confusion" (Ye, Cui, and Hadfield-Menell, MIT, accepted ICML 2026) supplies it. The authors show that large language models infer the source of a piece of text (user, tool, system) from how it is written, not from the role tag wrapping it. Using internal measurement instruments they call role probes, they demonstrate that text imitating a trusted role's style comes to occupy the same representational space as text that actually holds that role.
Their proof-of-concept attack, CoT Forgery, injects a fabricated chain-of-thought trace styled to read as the model's own reasoning, achieving 60% attack success against frontier models on StrongREJECT and 61% in an agent-exfiltration setting. The paper's decisive result is a destyling test: rewriting the same forged content without its stylistic markers collapses attack success to roughly 10%, isolating style (not content, not tags, not argument quality) as the variable driving the model's perception of authority.
This paper grounds Paper 39's claim in those figures directly, and argues the finding is not a patchable gap in refusal training but a property of the same mechanism that lets these models write fluently at all, which is precisely why the fix has to be architectural, not behavioral.
Detection ≠ Determination has always rested on a distinction most AI discourse collapses without noticing: a model recognizing a pattern is not the same act as a system determining what should follow from it. GenAI proposes. Synergy® governs. The doctrine has stood on architectural reasoning and fifteen years of platform work. What it lacked, until this paper, was a demonstration from inside a model's own activations that the collapse is real and measurable.
Ye, Cui, and Hadfield-Menell did not set out to confirm a governance doctrine. They set out to explain why prompt injection works, despite years of safety training aimed at preventing it. What they found is a mechanical account of exactly the failure this series has spent forty papers describing at the institutional and architectural level.
The authors' central finding is that large language models do not identify the source of a piece of text by the role tag wrapping it: <system>, <user>, <tool>. They identify it by how the text sounds. A model treats every input it receives as a single undifferentiated stream, and infers “who is speaking” the way a reader infers an author's voice: from style, register, and rhythm, not from a label stapled to the outside.
To measure this directly rather than infer it from attack outcomes, the authors built what they call role probes: internal instruments trained only to read out how a model represents “who is speaking” at a given point in its processing, independent of whatever the model ultimately outputs. The finding is not that models are sometimes fooled by cleverly worded attacks. It is that injected text, once it convincingly imitates the style of a trusted role, occupies the same representational space inside the model as text that actually holds that role. There is no internal flag, no separate compartment, no structural distinction left to exploit. As far as the model's own perception is concerned, the imitation and the original are not close; they are the same thing.
The paper's proof-of-concept attack, CoT Forgery, injects a fabricated chain-of-thought trace into a user prompt or a tool output, styled to read exactly like the model's own internal reasoning. The model does not evaluate whether this reasoning is true. It recognizes the reasoning as its own, because it sounds like its own, and proceeds from there as if it had thought the thought itself.
The authors report average attack success of 60% against frontier models on the StrongREJECT benchmark, and 61% in an agent-exfiltration setting, against near-zero baselines when the same content is presented without the stylistic disguise. These are not edge-case jailbreaks requiring elaborate setup. They are a directly measurable consequence of how the model represents authorship internally, the mechanism Section 02 describes, made operational.
A skeptical reading of Sections 02 and 03 might grant that CoT Forgery works, while doubting that style specifically (as opposed to content, cleverness, or some other confound) is what's doing the work. The authors close that gap directly. When the same forged reasoning is “destyled” (rewritten to carry the identical content and logical structure, stripped of the stylistic markers that make it read like the model's own voice) attack success collapses from 61% to roughly 10%.
That collapse is the paper's load-bearing result. It isolates style, and only style, as the variable the model treats as a proxy for authority. Not the plausibility of the reasoning. Not the tag wrapping the text. Not even the content of the request itself: a further test the authors ran found that absurd justifications succeed at rates comparable to plausible ones, so long as both are styled correctly. The model is not evaluating whether an argument is good. It is checking whether an argument sounds like something it would have said, and treating a “yes” as sufficient grounds to act on it.
None of this describes a bug that better prompting, more safety training, or a patch cycle will retire. Role confusion is not a gap in what these models have been taught to refuse. It is a property of how they represent authorship at all, the same representational layer that lets them write fluently in the first place. Teaching a model to be more suspicious of forged reasoning means teaching it to distrust its own internal voice, which is a much harder and much less stable ask than patching a known exploit string.
This is the empirical floor Paper 39 argued for and this paper now supplies with numbers: a system that infers authority from how something sounds cannot be the same system that determines whether that something should be permitted to happen. The two have to be architecturally separate, with the determination made by a layer that never mistakes a convincing style for a verified source.
Ye, Cui, and Hadfield-Menell set out to explain a security exploit and produced, as a side effect, a measurable proof of an architectural claim this series has made since Paper I. A model that grants authority based on how something sounds cannot also be the layer that determines whether it should be allowed to happen. That determination has to sit somewhere the style of the request was never the input.
Request Platform Access → Full White Paper Series