Role Without Determination

The Chain-of-Thought Forgery Problem

A paper accepted to ICML 2026 (the International Conference on Machine Learning, one of the field's most selective annual venues) found that large language models identify who is speaking not from the structural tags wrapped around a message, but from the style the text is written in. The resulting exploit, CoT Forgery, hijacks model behavior by writing fake reasoning in the model's own voice. This is Detection ≠ Determination, confirmed independently, from inside model internals.

Ken Granville CEO & Co-Founder, MindAptiv White Paper 39 The Governed Machine July 2026
Abstract

A paper titled "Prompt Injection as Role Confusion," authored by independent researchers Charles Ye and Jasmine Cui together with MIT associate professor Dylan Hadfield-Menell, was accepted to ICML 2026 and first covered publicly by MIT Technology Review on July 30, 2026, after earlier preprint coverage by security researcher Bruce Schneier and developer Simon Willison in June 2026.

The paper's finding, measured directly inside model internals using a technique the authors call role probes, is precise: large language models do not identify the source of a piece of text from the structural labels wrapped around it. They infer source from style. The resulting attack, CoT Forgery, plants fabricated reasoning styled to mimic a model's own chain-of-thought, and achieved roughly 60% attack success on a standard harmful-request benchmark across six frontier models, dropping to roughly 10% once the injected text's distinctive reasoning style was stripped out: empirical confirmation that style, not label, drives the vulnerability.

This paper argues the finding is independent, peer-reviewed validation of a claim this series has advanced since Paper I: Detection ≠ Determination. The researchers' own conclusion (that without a genuine mechanism for role perception, defending against this class of attack will remain "a perpetual whack-a-mole game") is the provenance fallacy (Paper XXXVIII) observed from inside a model's own representations.

Terms Used in This Paper
ICMLInternational Conference on Machine Learning: one of the AI field's most selective annual research conferences, where the paper discussed here was accepted.
LLMLarge Language Model: the general term for AI systems, such as ChatGPT, Claude, and Gemini, that generate text by predicting likely next words.
CoTChain-of-Thought: a scratch-pad process in which a model writes out step-by-step reasoning before producing a final answer.
CoT ForgeryThe attack this paper describes: fabricated reasoning, styled to mimic a model's own chain-of-thought, planted in its input to hijack its behavior.
Prompt InjectionAn attack in which unauthorized instructions are smuggled into a model's input, disguised as legitimate content from a trusted source.
Red-TeamingThe practice of deliberately attempting to break a system to find its weaknesses before an outside attacker does.
Role ProbesThe researchers' technique for inspecting a model's internal activations to measure how it represents "who is speaking" at a given point in a passage of text.

Section 01The Paper and What It Found

A paper titled "Prompt Injection as Role Confusion" (authored by independent researchers Charles Ye and Jasmine Cui together with MIT associate professor Dylan Hadfield-Menell) was accepted this year to ICML 2026, the International Conference on Machine Learning, one of the most selective annual research conferences in the AI field. The paper was first covered publicly by MIT Technology Review on July 30, 2026, and was picked up earlier in preprint form by security researcher Bruce Schneier and developer Simon Willison in June 2026.

The finding is precise and empirically measured. Large language models (LLMs, the general term for AI systems like ChatGPT, Claude, and Gemini that generate text by predicting likely next words) do not identify the source of a piece of text by the structural labels wrapped around it in a conversation (labels such as "this text is from the user," "this text is the system's own instructions," or "this text is the model's own internal reasoning"). Instead, models infer source from the style the text is written in. If a passage sounds like the model's own internal reasoning (commonly called chain-of-thought, or CoT, a scratch-pad process models use to work through a problem step by step before answering), the model tends to treat it as trustworthy, regardless of where it actually came from or what label was attached to it.

The researchers name this exploit CoT Forgery: an attacker plants fabricated reasoning text, styled to mimic a model's own internal chain-of-thought, inside a user prompt or inside content the model reads from elsewhere (a webpage, a document, a tool's output). The model then treats that forged reasoning as though it had produced it itself, and acts on its fabricated conclusion.

The Measured Result
Tested against six frontier models with no special access and no iterative tuning of the attack, CoT Forgery achieved roughly 60% attack success on a standard harmful-request benchmark, compared to near-zero for the same requests presented without forged reasoning. Stripping the distinctive "reasoning voice" out of the injected text while leaving its content intact dropped success from roughly 61% to about 10%: confirmation that style, not the structural label, is what the model is actually keying off of.

Section 02Why This Is a Doctrinal Confirmation, Not a New Discovery

MindAptiv's founding position, Detection ≠ Determination, has held since the company's earliest public materials that pattern recognition of an input (what something looks like) is categorically distinct from a governed determination of what that input is authorized to do. GenAI proposes; Synergy® governs. The distinction exists precisely because proposal-generating systems are, by construction, style-matchers: they are extraordinarily good at recognizing what a system instruction, a user message, or a reasoning trace typically looks like, and they have no independent mechanism for verifying that a given block of text is what it claims to be.

The ICML paper describes this exact failure mode from the inside, using role probes, tools that let the researchers look inside a model's internal activations and measure how the model is representing "who is speaking" at each point in a passage of text. What they found is that forged text lands in the same internal representational space as the genuine role it's imitating. The model isn't being tricked by a mislabeled tag; it never was reading the tag as authoritative in the first place. It was always inferring role from style, and an attacker who writes in the right style is, from the model's internal point of view, indistinguishable from the real thing.

This is not a wolf whose disguised voice is nearly good enough to pass. It is a door with no peephole at all, answering on the knock and the claim alone.

The Confirmation, Stated Plainly
This is Detection operating unaccompanied by Determination, measured directly inside model internals, by researchers with no relationship to MindAptiv or to the Wantware paradigm, arriving independently at the same structural diagnosis this series has argued from first principles since Paper I.

Section 03The Unsolvability Claim

The more consequential claim in the paper is not that this particular exploit exists, but that the researchers regard the underlying problem as likely permanent absent a genuinely different approach to how models represent roles. The authors state plainly that without a real mechanism for role perception, defending against prompt injection (the broader category of attack in which someone smuggles unauthorized instructions into a model's input) will remain "a perpetual whack-a-mole game." Their reasoning: red-teaming (the practice of deliberately trying to break a system to find its weaknesses before an attacker does) and adversarial training can only ever produce a list of known bad patterns to refuse, and no such list is exhaustive. Every new list trains the model to resist a class of attacks that already worked, while doing nothing structural to prevent the next stylistic variant.

Cross-reference · This is consistent with what this series has termed the provenance fallacy (Paper XXXVIII): the assumption that a system can be made trustworthy by getting better at recognizing untrustworthy inputs, when the actual gap is that the system has no independent, non-linguistic mechanism for establishing provenance in the first place.

Recognition improves. Determination does not exist. The gap does not close. It just gets harder to find.

Section 04Relevance to the Essence Architecture

The MindAptiv Essence® platform's separation of concerns is directly responsive to the failure mode this research describes. Three platform components are relevant.

SecuriSync™
Trust Before Execution
SecuriSync™ does not ask whether a given instruction looks like an authorized instruction. It is a platform-wide identity and trust substrate (see also the Cybersecurity overview) that renders a governed decision about whether execution is permitted at all, evaluated before the action runs rather than inferred from the model's own token stream.
Aptivs
Declared Purpose, Verified Meaning
Aptivs, the core portable, composable unit of the Essence platform, each carrying its own declared purpose, verified meaning, and allowed outcomes, do not rely on a model's internal sense of "who said this" to determine what they are permitted to do. Every Aptiv is bound to a Guard, described on the Nebulo page, which governs access, trust, and behavior at a granular level as a structural property of the Aptiv itself.
Guard / Nebulo
Structural, Not Self-Reported
Trust level is structural, encoded in how an Aptiv is built and what it is permitted to touch, not a policy the model is trusted to remember and re-apply correctly to every new stylistic disguise an instruction might wear.
How Essence® Resolves It
SecuriSync decides if you can run.
Guard ensures you behave while running.
Neither check depends on the model's chain-of-thought being genuine. The model's chain-of-thought is not the authority on what the model is permitted to do next.

Put directly: CoT Forgery works by getting a model to trust its own text. An architecture that never asked the model to be the arbiter of its own provenance in the first place has no equivalent point of failure to exploit, not because it is harder to fool, but because "fooling the model's self-assessment" ceases to be a meaningful attack surface when the model's self-assessment was never the thing making the determination.

Section 05The Governed Machine's Position

Independent, peer-reviewed research now describes, from inside the internal representations of transformer-based models, the exact structural gap this series has named since its founding papers. Detection, however refined, is a statement about what an input resembles. Determination is a statement, made by something other than the pattern-matcher itself, about what that input is authorized to do. The ICML findings suggest this gap is not a training deficiency awaiting a fix. It is architectural.

Systems that continue to ask a language model to adjudicate its own inputs will keep finding new variants of this exploit indefinitely. Systems that never delegated that adjudication to the model in the first place have already closed the door the researchers are describing.

Primary source · Ye, C., Cui, J., & Hadfield-Menell, D. (2026). Prompt Injection as Role Confusion. Accepted to ICML 2026. Project page: role-confusion.github.io · ICML listing: icml.cc/virtual/2026/poster/64605 · Secondary coverage: MIT Technology Review, July 30, 2026 · The Register, June 30, 2026 · Hackaday, July 2, 2026
A Note on Scope
MIT Technology Review reports, secondhand from the researchers, that similar results have since been observed against models made by Anthropic, Alibaba, and DeepSeek, beyond the OpenAI-model results detailed in the paper itself. No methodology for those follow-on tests is described in that article; this should be treated as a reported claim pending review of any published follow-up. Neither OpenAI nor Anthropic is reported to have corroborated the specific attack transcripts described in press coverage.
The Governed Machine: Paper 39

The model's chain-of-thought
was never the authority.

Independent, peer-reviewed research now describes, from inside the internal representations of transformer-based models, the exact structural gap this series has named since its founding papers. Detection, however refined, is a statement about what an input resembles. Determination is a statement, made by something other than the pattern-matcher itself, about what that input is authorized to do. Systems that continue to ask a language model to adjudicate its own inputs will keep finding new variants of this exploit indefinitely.

Request Platform Access → Full White Paper Series

White Paper Series · The Governed Machine

1The Civilizational Fault Line 2We Are Building the Wrong Machine 3The Ornithopter Mistake 4The Convergence 5The Four Horsemen of the Knowledge Apocalypse 6What the Insiders Confirmed 7The Metaphor Trap 8The Recall Standard 9The $1 Trillion Governance Gap 10The Litigation Layer 11The Scale of Intent 12The Intent Economy 13The Session Illusion 14The Necessary Sequence 15The Wrong Race 16The Ledger That Is Intent-Driven 17The Agency Illusion 18The Substrate 19The End of the Mean 20Era 3: The Architecture of the Next Civilization 21The Missing Substrate 22The Context Fatigue Ceiling 23The Iceberg Stays Frozen 24The Dependency Tax 25The Record That Was Never Kept 26Composable by Default 27Do No Harm 28The Stack Replacement Thesis 29The Moat Is the Code 30The Last Platform War 31Beyond the Agent: Intent-Native Execution 32The Hardware Imagination 33The Architecture Tax 34The Tokenization Ceiling 35The Payment Moment 36The Oracle Problem 37The Reviewer Problem 38The Provenance Fallacy 39Role Without Determination ← this paper 40Known and Funded Anyway 41The Style Confusion Proof 42The Verification Tax 43The Pause Reflex 44The Human Margin 45The Balance of Power Fallacy 46The Liability Backstop 47One Substrate, Every Signal 48The Attribution Problem 49The Consciousness Ceiling 50The Detection Patch 51The Consumptive Machine 52The Agent That Isn't 53The Legibility Gap 54The Semiotic Machine 55The Transpilation Ceiling 56The Provisioning Ceiling 57The Reservation Ceiling 58The Circularity Ceiling 59The Coexistence Ceiling 60The Conformance Ceiling 61The Preservation Ceiling 62The Parity Clause 63The Governed Boundary 64The Transcript Problem 65The Unpaired System 66The Memory Ceiling 67The Admission Gap 68The Wrong Ask 69The Best Case 70The Last Chokepoint 71The Fourth Step 72The Adoption Standard 73The Same Weekend 74Sixty to One 75Coordinates, Not Correlations 76The Governability Axis 77Era 3, Confirmed 78The Eleventh Rule 79The Seventh Admission 80The Authorization Gap 81The Authorship Fallacy 82The Camera and the Vault 83Cleared to Proceed 84A Class, Not a Product 85The Inherited Playbook