The Chain-of-Thought Forgery Problem
A paper accepted to ICML 2026 (the International Conference on Machine Learning, one of the field's most selective annual venues) found that large language models identify who is speaking not from the structural tags wrapped around a message, but from the style the text is written in. The resulting exploit, CoT Forgery, hijacks model behavior by writing fake reasoning in the model's own voice. This is Detection ≠ Determination, confirmed independently, from inside model internals.
A paper titled "Prompt Injection as Role Confusion," authored by independent researchers Charles Ye and Jasmine Cui together with MIT associate professor Dylan Hadfield-Menell, was accepted to ICML 2026 and first covered publicly by MIT Technology Review on July 30, 2026, after earlier preprint coverage by security researcher Bruce Schneier and developer Simon Willison in June 2026.
The paper's finding, measured directly inside model internals using a technique the authors call role probes, is precise: large language models do not identify the source of a piece of text from the structural labels wrapped around it. They infer source from style. The resulting attack, CoT Forgery, plants fabricated reasoning styled to mimic a model's own chain-of-thought, and achieved roughly 60% attack success on a standard harmful-request benchmark across six frontier models, dropping to roughly 10% once the injected text's distinctive reasoning style was stripped out: empirical confirmation that style, not label, drives the vulnerability.
This paper argues the finding is independent, peer-reviewed validation of a claim this series has advanced since Paper I: Detection ≠ Determination. The researchers' own conclusion (that without a genuine mechanism for role perception, defending against this class of attack will remain "a perpetual whack-a-mole game") is the provenance fallacy (Paper XXXVIII) observed from inside a model's own representations.
A paper titled "Prompt Injection as Role Confusion" (authored by independent researchers Charles Ye and Jasmine Cui together with MIT associate professor Dylan Hadfield-Menell) was accepted this year to ICML 2026, the International Conference on Machine Learning, one of the most selective annual research conferences in the AI field. The paper was first covered publicly by MIT Technology Review on July 30, 2026, and was picked up earlier in preprint form by security researcher Bruce Schneier and developer Simon Willison in June 2026.
The finding is precise and empirically measured. Large language models (LLMs, the general term for AI systems like ChatGPT, Claude, and Gemini that generate text by predicting likely next words) do not identify the source of a piece of text by the structural labels wrapped around it in a conversation (labels such as "this text is from the user," "this text is the system's own instructions," or "this text is the model's own internal reasoning"). Instead, models infer source from the style the text is written in. If a passage sounds like the model's own internal reasoning (commonly called chain-of-thought, or CoT, a scratch-pad process models use to work through a problem step by step before answering), the model tends to treat it as trustworthy, regardless of where it actually came from or what label was attached to it.
The researchers name this exploit CoT Forgery: an attacker plants fabricated reasoning text, styled to mimic a model's own internal chain-of-thought, inside a user prompt or inside content the model reads from elsewhere (a webpage, a document, a tool's output). The model then treats that forged reasoning as though it had produced it itself, and acts on its fabricated conclusion.
MindAptiv's founding position, Detection ≠ Determination, has held since the company's earliest public materials that pattern recognition of an input (what something looks like) is categorically distinct from a governed determination of what that input is authorized to do. GenAI proposes; Synergy® governs. The distinction exists precisely because proposal-generating systems are, by construction, style-matchers: they are extraordinarily good at recognizing what a system instruction, a user message, or a reasoning trace typically looks like, and they have no independent mechanism for verifying that a given block of text is what it claims to be.
The ICML paper describes this exact failure mode from the inside, using role probes, tools that let the researchers look inside a model's internal activations and measure how the model is representing "who is speaking" at each point in a passage of text. What they found is that forged text lands in the same internal representational space as the genuine role it's imitating. The model isn't being tricked by a mislabeled tag; it never was reading the tag as authoritative in the first place. It was always inferring role from style, and an attacker who writes in the right style is, from the model's internal point of view, indistinguishable from the real thing.
This is not a wolf whose disguised voice is nearly good enough to pass. It is a door with no peephole at all, answering on the knock and the claim alone.
The more consequential claim in the paper is not that this particular exploit exists, but that the researchers regard the underlying problem as likely permanent absent a genuinely different approach to how models represent roles. The authors state plainly that without a real mechanism for role perception, defending against prompt injection (the broader category of attack in which someone smuggles unauthorized instructions into a model's input) will remain "a perpetual whack-a-mole game." Their reasoning: red-teaming (the practice of deliberately trying to break a system to find its weaknesses before an attacker does) and adversarial training can only ever produce a list of known bad patterns to refuse, and no such list is exhaustive. Every new list trains the model to resist a class of attacks that already worked, while doing nothing structural to prevent the next stylistic variant.
Recognition improves. Determination does not exist. The gap does not close. It just gets harder to find.
The MindAptiv Essence® platform's separation of concerns is directly responsive to the failure mode this research describes. Three platform components are relevant.
Put directly: CoT Forgery works by getting a model to trust its own text. An architecture that never asked the model to be the arbiter of its own provenance in the first place has no equivalent point of failure to exploit, not because it is harder to fool, but because "fooling the model's self-assessment" ceases to be a meaningful attack surface when the model's self-assessment was never the thing making the determination.
Independent, peer-reviewed research now describes, from inside the internal representations of transformer-based models, the exact structural gap this series has named since its founding papers. Detection, however refined, is a statement about what an input resembles. Determination is a statement, made by something other than the pattern-matcher itself, about what that input is authorized to do. The ICML findings suggest this gap is not a training deficiency awaiting a fix. It is architectural.
Systems that continue to ask a language model to adjudicate its own inputs will keep finding new variants of this exploit indefinitely. Systems that never delegated that adjudication to the model in the first place have already closed the door the researchers are describing.
Independent, peer-reviewed research now describes, from inside the internal representations of transformer-based models, the exact structural gap this series has named since its founding papers. Detection, however refined, is a statement about what an input resembles. Determination is a statement, made by something other than the pattern-matcher itself, about what that input is authorized to do. Systems that continue to ask a language model to adjudicate its own inputs will keep finding new variants of this exploit indefinitely.
Request Platform Access → Full White Paper Series