Coordinates, Not Correlations

Why Attention Was Never Built to Answer What AI Governance Requires

In 2017, eight researchers at Google showed that attention alone, with no recurrence and no convolution, could translate language better and train faster than anything before it. That result became the substrate under nearly every frontier model. Attention was never built to say what a task actually meant. Asking a correlation to do a coordinate's job explains why the debate over how to govern AI keeps missing its most critical point: a check on each action, against something fixed, before it runs. Without coordinates, interpretability can show what a model is doing but cannot determine whether an action is authorized.

Ken Granville CEO & Co-Founder, MindAptiv White Paper 75 The Governed Machine September 2026
Abstract

"Attention Is All You Need," published by Vaswani et al. in 2017, replaced recurrence and convolution with a single mechanism, self-attention, and in doing so removed the sequential bottleneck that had made prior models slow to train. The resulting Transformer architecture trained faster, translated better, and generalized to tasks well outside its original brief. It is, without qualification, the correct foundation beneath most current frontier models. This paper credits that achievement in full, then draws a precise distinction the achievement itself does not resolve. Attention answers which tokens in a given sequence are relevant to one another, a relationship learned from training data and re-derived at every inference pass. It has never been built to answer what a specific task was asked to accomplish, independent of the sequence that happens to describe it. That distinction, correlation versus coordinate (what MindAptiv calls a Meaning Coordinate), has gone unnamed: the industry's remedies stop at three steps, embedded evaluators, certification checkpoints and coordination, all detection or pacing. The fourth step, an action-level determination check, has stayed unclaimed not for lack of proposals but because nothing upstream in a correlation-based architecture produces a fixed thing to check a proposed action against. The error is not attention. It is asking a correlation to do a coordinate's job.

Section 01The Paper Examined at Its Best

Before 2017, the dominant approach to sequence tasks such as translation ran recurrent or convolutional networks in an encoder-decoder configuration, generating hidden states one position at a time. That sequential structure blocked parallelization within a training example, which became a hard ceiling as sequence length grew. Vaswani and seven co-authors proposed the Transformer: a network built solely on attention mechanisms, with no recurrence and no convolution. The result trained in a small fraction of the time the strongest prior models required, reached a new best single-model BLEU score on the WMT 2014 English-to-French translation task, and improved on the best prior ensemble results on the English-to-German task. It also generalized cleanly to English constituency parsing, a task well outside translation.

That is not an incremental result. It is the correct architectural foundation beneath most current frontier models, agentic and otherwise, and it deserves the same treatment given to Falcon Guardian in Paper 74: credited fully, on its own terms, before any distinction is drawn against it.

Section 02What Attention Actually Finds

Self-attention computes, for each element of a sequence, a weighted relevance to every other element, using representations learned entirely from training data. Multiple attention heads let the model track several such relationships at once. This is a genuine advance over fixed-window or strictly sequential relevance, and it is why a single word can be resolved differently depending on the sentence around it. But every part of that mechanism is contextual and re-derived: the weights come from the specific sequence in front of the model, at the specific pass being computed, using representations shaped by whatever the model was trained on. Change the surrounding tokens, the model checkpoint, or the training corpus, and the same input can resolve to a different attended relationship.

That is a correlation, in the precise sense used here: a statistically inferred association between things, recomputed each time it is asked for, with no persistent reference outside the computation that produced it.

Section 03Two Different Questions

Two Different Questions Being Answered
What Attention Asks
Given this token, in this sequence, which other tokens should it attend to, weighted by relevance learned from training data?
What a Coordinate Answers
What is the fixed destination this content maps to, independent of the specific sequence, model checkpoint, or inference pass that happens to encounter it?

Neither question is the wrong one to ask. They are different questions, built to do different work. Attention's question produces excellent translations, coherent completions, and useful agentic behavior across an enormous range of tasks, precisely because most of those tasks reward a good contextual approximation. A coordinate's question produces something else: a reference that stays the same regardless of which model, which pass, or which surrounding sequence is asking. Only one of those two things can be checked against.

A destination reached by air, train, car, bike, or on foot does not change. Routes do. They can be fastest or slowest, most scenic or most dangerous, and each starts somewhere different and demands different decisions. They all still end up at the same place: the coordinate. Even the dangerous route is useful, because with a fixed destination a decision to avoid it can be made.

Section 04The Gap Correlation Can't Close

An attention-weighted representation of "what this means" is a route, not a destination. It gets recomputed per pass, shifts with context length and phrasing, and carries no persistent form that a later, independent check could compare against. For generation and translation, that is not a defect: a good-enough approximation, freshly computed each time, is the entire point. But determination is a different task from generation. Determining whether a specific proposed action matches what a task actually called for requires something to check the action against that does not change depending on which pass produced the action being checked. A moving target cannot serve as its own reference.

Two Primitives, Compared on Stated Axes
AxisAttentionCoordinate
Computed fromLearned relevance across tokens in the current sequenceA fixed primitive in a published, closed set
Persists across inference passesNo: recomputed every passYes: the same value regardless of pass
Depends on training corpusYes: shifts if the model or its training data changesNo: independent of any model or training run
Stable under paraphraseApproximate: similar phrasing usually resolves similarly, not guaranteedYes: any phrasing with a match resolves to the same coordinate; a new phrasing gets one when someone explains what it means
Checkable by an independent systemNo: no persistent form outside the computation that produced itYes: a separate system can compare a proposed action against the same fixed reference

This is not a claim that attention-based models are unreliable at what they were built to do. It is a claim that a fixed, checkable answer to what a task actually meant was not what attention was designed to produce, no matter how much larger the model or how much more training data it sees.

A Meaning Coordinate is one of 256 fixed primitives, grouped into four realms (Operations, Cognition, Physical, Relational) and published as a closed set at mindaptiv.com/meaning-coordinates. An intent resolves to a short combination of them. The codes in the worked example below are shown only to make that concrete; Section 07 covers how they are produced and used.

No one writes the coordinate column by hand. A person gives the instruction in plain language; the resolution to primitives happens underneath it, the same way a chemist's formula describes what a cake's ingredients are doing without the baker ever needing to see it. The coordinates are shown here only to make the distinction concrete, not because producing them is something a user, or even most engineers, would ever do directly.

A Worked Example · "Detect anomalous transaction pattern, freeze account, and generate audit trail"
PhraseCoordinateMeaning
Detect anomalous transaction patternMz·R4 BryKrz·R4 ZeDe·R1Sense (the emitted signal) + Formula·Compare (rule-based pattern detection) + Difference·Default (deviation from the declared normal)
freeze accountSwe·R3 DiKo·R2Snare (lock the entity, freeze its state) + Thing·System (the named account entity)
and generate audit trailChzJe·R1 TuSe·R2 Che·R1Create·Chronicle (initiate the event record) + Relation·Target (append to the governed ledger) + Guard (tamper-proof enforcement)

Rephrase that instruction ("flag the suspicious transfer, lock the account, log what happened") and an attention-based model may resolve it to a somewhat different internal weighting, since the words and their surrounding context have changed. The coordinates above do not move, because they are fixed primitives, not a weighting recomputed from the words used. What connects different wording to them is a match, and a phrase with no match yet is given one by explaining what it means. That is the entire distinction this paper is making, made concrete: not that one resolution is smarter than the other, but that only one of them stays the same when the sentence describing it does not.

Section 05Why the Fourth Step Stayed Unclaimed

The error is not attention. It is asking a correlation to do a coordinate's job: treating a better-trained route as if it were a fixed destination.

Papers 68 through 74 examined a succession of proposals for checking an agent's action before or as it executes: capability checkpoints, embedded evaluators, allow lists, signature matching against known-bad behavior. Each was evaluated on its own terms, and each was found to fall into one of two categories, detection or pacing, never the fourth step, defined as a determination check performed on a specific proposed action, at the moment it is proposed, independent of the model's disposition or its last certification.

What none of those papers named directly is why every proposal kept landing in the same two categories. A determination check needs a fixed thing to determine against. An attention-based system has no such thing anywhere in its architecture: what it has, at every layer, is a re-derived correlation, useful for generating the next output and unsuitable as a stable reference for judging a later one. A system with no fixed destination cannot distinguish a wrong turn from a right one; every action looks equally plausible once nothing exists to weigh it against. The fourth step did not stay unclaimed because nobody thought to propose it. It stayed unclaimed because nothing upstream in the dominant architecture produces the one ingredient a determination check requires.

See also: Paper 71, "The Fourth Step": on Amodei's own concession that capability checkpoints and interpretability are detection performed earlier, not determination, and Paper 74, "Sixty to One": on why even excellent signature-based enforcement cannot catch an approved agent taking an unauthorized action that matches no known pattern.

Section 06Why It Can't Be Bolted On

Governance layered onto a correlation-based system after the fact has nothing native inside that system to attach to. There is no fixed coordinate sitting inside a Transformer's forward pass that a governance layer could simply read and check; there is only a distribution of attention weights, valid for one pass, gone by the next. Every governance approach examined here, from capability checkpoints to Falcon Guardian's allow lists, is consequently built the only way it can be: as an external wrapper watching inputs and outputs, rather than a determination performed natively as part of the execution it is supposed to govern.

A coordinate-based system does not face that constraint, because the fixed reference is not added afterward; it is what the primitive already is. Checking an action against a coordinate is not a separate department bolted onto the architecture, in the same way that a location on a map does not require a separate department to confirm it is the correct location; correctness is the coordinate's native property. A correlation has no equivalent property to extend. This is the deepest answer available to the question raised implicitly above: why does every remedy examined so far look like an addition rather than a foundation. It is not a failure of effort, urgency, or resourcing on the part of any lab or vendor examined. It is an architectural fact: a destination cannot be retrofitted onto a system that was only ever built to compute routes.

This is also why interpretability, however much it reveals, remains detection. What it recovers are representations specific to one model, one checkpoint and one training run, and an action cannot be checked against something that changes whenever the model does. Understanding why a model acts as it does is not the same as holding a fixed reference for whether a specific action is authorized.

Section 07The Doctrine, One Level Lower

The Doctrine, Read One Level Lower
Detection ≠ Determination. Correlation ≠ Coordinate. Code ≠ Intent.
Three layers, one absence. Security tooling reconstructs authorization after an action runs. A correlation-based model reconstructs relevance after each pass. An agent reconstructs the intent behind its own code from logs, after the fact. Papers 11 and 71 through 74 name the same missing thing at three different altitudes.
The MindAptiv Position
This is precisely what Meaning Coordinates are built to supply: a closed, published set of primitives, functioning the way a location on a map functions, that a proposed action can be checked against directly rather than approximated toward. Not a larger attention window, and not a better-trained correlation, both of which improve the quality of the route, but a fixed destination the route can finally be checked against. Essence®'s propose/determine/execute architecture exists to perform that check natively, at the moment an action is proposed, using a reference that does not change depending on which pass produced the proposal. The coordinate itself does not execute; it is the fixed reference an action is checked against. What executes is the machine instruction synthesized from it, and that synthesis is free to adapt continuously to the hardware it runs on without the coordinate underneath it ever needing to move.

It is worth tracing the full path from a plain-language request to a running action, because GenAI has a real role in it, and what carries out the resulting action is not what is usually called an agent.

Intent to Execution
Intent: Natural Language
A person states what they want in plain language. Informal, ambiguous, unresolved: the same kind of input an attention-based system would also take in.
GenAI Proposes
A model interprets the request and proposes a candidate resolution, drawing on the reasoning and synthesis GenAI is genuinely good at. The proposal is not yet authorized to do anything.
Meaning Coordinates
Synergy® resolves the proposal into fixed primitives from the published 256-coordinate set. This is the reference the resulting action will be checked against, addressed directly rather than approximated from wording.
PowerAptiv Persists · Executes
The governed record materializes as a Trust-Certified PowerAptiv (the execution-type Aptiv among Essence®'s eight Aptiv types) anchored to the coordinates and their provenance. It converts intent into optimized machine instructions in real time, or chaperones existing code inside the same governed flow, not as an autonomous agent taking probabilistic steps, but as a bounded action authorized before it runs.

That last distinction is not cosmetic. Elsewhere, "agent" describes a system that acts on a learned, re-derived sense of what to do next, which is exactly the correlation problem Sections 02 through 06 describe. A PowerAptiv is materialized from a fixed coordinate envelope and cannot produce an action outside it, not because a monitor is watching for violations, but because the action was never encoded into what the PowerAptiv is.

See also: Paper 11, "The Scale of Intent": on why an agent's computational primitive is code, requiring its authorization to be reconstructed after the fact from execution logs, while an Aptiv's primitive is intent, carrying its authorization natively and persistently. That is the same absence this paper names one layer down: an agent reconstructs intent after code runs, the way attention reconstructs relevance after each pass.

How the resolution step earns its authority matters, because a fixed coordinate is only as good as the mapping that produces it. In Essence®, that mapping is not a model. Synergy® matches the natural-language proposal against Grok Units, structured templates authored by people, each of which resolves a pattern of natural language to specific coordinates. The same input always resolves to the same coordinates, and each resolution is recorded with the input, the coordinates and the rules applied. That guarantees reproducibility and auditability, not infallibility: a template that is wrong will be reproducibly wrong. This is why the authority sits with the people who author and approve the templates, and not with a statistical estimate of what was meant.

When a phrase has no match, the system does not infer a meaning. A person explains what it means, in real time, and the match is created. A bank team that says "put it on ice" to mean freezing an account can tell Synergy® so, within the context of its own operations, and from then on that phrase resolves to the same coordinates as "freeze account" in the worked example in Section 04. The addition is scoped to that context and is additive: it replaces nothing, and a team that uses the same phrase differently is unaffected unless it approves the new match. How matches are shared across individuals and organizations depends on their own policies and requirements, and is outside the scope of this paper. The explanation is the authority: the mapping exists because a person said what the phrase means.

One Layer Fixed, One Layer Free to Adapt
The Coordinate: Fixed
A published primitive from the 256-coordinate set. Does not execute. Does not change with hardware, phrasing, or which model produced the proposal. This is the reference a proposed action is checked against.
The Execution: Adaptive
Machine instructions synthesized from the coordinate, continuously, against the hardware actually present at that moment. At Trust Level 3, a PowerAptiv is codeless (pure Meaning Coordinates) and its instruction synthesis is never the same twice: there is no fixed artifact underneath it to extract or reverse. This layer is supposed to change as conditions change; the coordinate above it does not.
Meaning Coordinates · 256 Primitives, Four Realms
RealmWhat It Covers
R1 · OperationsThe acts of comparing, collecting, counting, and measuring that underpin all computation
R2 · CognitionValues, logic, information, thought, instances, situations, possibilities, and work
R3 · PhysicalSpacetime, physical properties, scale, shape, energy, matter, body, and thing types
R4 · RelationalSharing, perceiving, judging, feeling, moving, doing, using, and communicating
The full set (all 256 primitives across 32 groups, with the live coordinate map) is published at mindaptiv.com/meaning-coordinates.

Section 08What Changes and What Doesn't

This does not change the argument. It supplies the mechanism underneath it. Attention Is All You Need earned its place as the substrate of modern AI by solving a real and hard problem: how to relate every part of a sequence to every other part, quickly and in parallel. It was never posed as a solution to a different problem, fixing what a task was actually asked to accomplish, and nothing in the nine years since has required it to become one. The fourth step, named in Paper 71, was never missing because of insufficient effort at the evaluation or enforcement layer. It was missing because the layer beneath all of them was built to find routes, not destinations.

What should change, for anyone evaluating whether a bigger model, a better evaluator, or a more complete signature library will eventually close this gap, is the recognition that none of those are the right kind of thing to add. A coordinate is a different kind of primitive than a correlation, however well-trained. Until one exists inside the architecture being governed, the fourth step remains exactly what it was before this paper: unclaimed, and unclaimable by addition alone.

Section 09References

External sources

  1. [1]Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). “Attention Is All You Need.” Advances in Neural Information Processing Systems 30 (NIPS 2017). arXiv:1706.03762 · NeurIPS proceedings
  2. [2]Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). “Bleu: a Method for Automatic Evaluation of Machine Translation.” Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL). ACL Anthology
  3. [3]Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., Soricut, R., Specia, L., and Tamchyna, A. (2014). “Findings of the 2014 Workshop on Statistical Machine Translation.” Proceedings of the Ninth Workshop on Statistical Machine Translation (WMT 2014). ACL Anthology · WMT 2014 translation task

Papers in this series cited above

  1. Paper 11The Scale of Intent
  2. Paper 67The Admission Gap
  3. Paper 68The Wrong Ask
  4. Paper 69The Best Case
  5. Paper 70The Last Chokepoint
  6. Paper 71The Fourth Step
  7. Paper 72The Adoption Standard
  8. Paper 73The Same Weekend
  9. Paper 74Sixty to One
The Governed Machine: Paper 75
Attention finds routes.
It was never built to find destinations.
Attention Is All You Need solved a real problem and deserves full credit for solving it. But every layer of it is a correlation, re-derived per pass, with nothing fixed to check a later action against. That is why the fourth step has stayed unclaimed, and why it cannot be bolted on from outside.
Request Access Read Paper 74

White Paper Series · The Governed Machine

1The Civilizational Fault Line 2We Are Building the Wrong Machine 3The Ornithopter Mistake 4The Convergence 5The Four Horsemen of the Knowledge Apocalypse 6What the Insiders Confirmed 7The Metaphor Trap 8The Recall Standard 9The $1 Trillion Governance Gap 10The Litigation Layer 11The Scale of Intent 12The Intent Economy 13The Session Illusion 14The Necessary Sequence 15The Wrong Race 16The Ledger That Is Intent-Driven 17The Agency Illusion 18The Substrate 19The End of the Mean 20Era 3: The Architecture of the Next Civilization 21The Missing Substrate 22The Context Fatigue Ceiling 23The Iceberg Stays Frozen 24The Dependency Tax 25The Record That Was Never Kept 26Composable by Default 27Do No Harm 28The Stack Replacement Thesis 29The Moat Is the Code 30The Last Platform War 31Beyond the Agent: Intent-Native Execution 32The Hardware Imagination 33The Architecture Tax 34The Tokenization Ceiling 35The Payment Moment 36The Oracle Problem 37The Reviewer Problem 38The Provenance Fallacy 39Role Without Determination 40Known and Funded Anyway 41The Style Confusion Proof 42The Verification Tax 43The Pause Reflex 44The Human Margin 45The Balance of Power Fallacy 46The Liability Backstop 47One Substrate, Every Signal 48The Attribution Problem 49The Consciousness Ceiling 50The Detection Patch 51The Consumptive Machine 52The Agent That Isn't 53The Legibility Gap 54The Semiotic Machine 55The Transpilation Ceiling 56The Provisioning Ceiling 57The Reservation Ceiling 58The Circularity Ceiling 59The Coexistence Ceiling 60The Conformance Ceiling 61The Preservation Ceiling 62The Parity Clause 63The Governed Boundary 64The Transcript Problem 65The Unpaired System 66The Memory Ceiling 67The Admission Gap 68The Wrong Ask 69The Best Case 70The Last Chokepoint 71The Fourth Step 72The Adoption Standard 73The Same Weekend 74Sixty to One 75Coordinates, Not Correlations ← this paper 76The Governability Axis 77Era 3, Confirmed 78The Eleventh Rule 79The Seventh Admission 80The Authorization Gap 81The Authorship Fallacy 82The Camera and the Vault 83Cleared to Proceed 84A Class, Not a Product 85The Inherited Playbook