What Sutton and LeCun Get Right About World Models, and How Essence Builds One Differently
Two Turing Award winners agree that large language models are not the whole of intelligence, and have staked their careers, in Rich Sutton's case, and a $3.5 billion company, in Yann LeCun's, on saying so publicly. Neither is arguing about tokens for its own sake. Both are arguing that intelligence requires a real model of the world, one built from experience, the other from self-supervised prediction. Semiotics named the underlying split a century before either of them did. This paper uses that vocabulary to trace both critiques to the same requirement, then shows what a third way of meeting it, built to be declared rather than learned, looks like next to both of them.
In a recent, widely covered interview about his Oak Lab's work, Turing Award winner Rich Sutton restated a critique he has been making for some time: language may account for only 20 to 25 percent of intelligence, and an industry that treats large language models as a stand-in for the whole of AI is committing a category error. Sutton is not alone in that view. Yann LeCun, a 2018 Turing Award recipient and Meta's chief AI scientist until his departure in November 2025, has built an entire company, AMI Labs, on the argument that large language models are a dead end for real intelligence, and that what is missing is a world model. Taken on its substance, neither critique is new to every field, only to this one. Semiotics, the study of signs, formalized the distinction between a sign and what it signifies more than a hundred years before the transformer architecture existed, in the work of Ferdinand de Saussure and Charles Sanders Peirce.
This paper argues that Sutton's language-versus-intelligence gap, LeCun's token-versus-world-model gap, and the semiotic sign-versus-meaning gap are the same structural observation viewed from three different disciplines. Where this paper departs from treating that as an intelligence problem to be solved is the goal it assigns to Essence. Essence was not built to make machines more intelligent, and it does not compete with Sutton's continual-learning agenda, LeCun's self-supervised representation learning, or a language model's fluency. It was built on the premise that a human's intent, and a governed model of the world that intent acts on, should reach a machine as declared, structural facts, not as signals a machine has to interpret its way toward. Sutton and LeCun are both evidence for why that premise matters, not targets this paper is trying to out-compute.
Rich Sutton is not a critic on the outside of the field looking in. He is a 2024 Turing Award recipient, one of the two people most associated with founding reinforcement learning as a discipline, and the co-founder, with Khurram Javed, of Oak Lab, which is building toward continual learning systems under what the lab calls the Big World Hypothesis: the premise that the world is too complex for any single agent to model completely, so intelligence has to come from an agent that keeps learning from its own experience rather than one trained once on a frozen corpus and then deployed. In that context, his comments about language models read less like a dismissal and more like a specific, bounded claim. Language, he argues, is a real and impressive capability, and large language models represent a genuine scientific breakthrough. What frustrates him, in his own framing, is that this progress in a subset of AI is being asked to stand in for the whole of it, when language itself is only a fraction, by his estimate 20 to 25 percent, of what intelligence actually requires.
Paper XXXIV, The Tokenization Ceiling, made a related point from a different direction: an architecture built around predicting the next token inherits a ceiling from that choice no matter how much compute is thrown at it, because the token is a proxy for meaning, not meaning itself. Sutton is making a claim about learning and goal-directed behavior. This series has been making a claim about representation. They are not the same claim, but the useful thing to take from Sutton is not that this series has a better answer to intelligence. It is that he is confirming, from reinforcement learning, that the entire premise of getting a machine to guess correctly what a human meant, however well it guesses, is a different project from letting the human specify what they meant directly and having the machine execute it without a guess in the middle. Essence was built for the second project. It has never been trying to win the first one.
Sutton is not the only Turing Award winner who has recently staked a career, and a great deal of capital, on the claim that large language models are not the path to real machine intelligence. Yann LeCun, who shared the 2018 Turing Award with Yoshua Bengio and Geoffrey Hinton and spent twelve years as Meta's chief AI scientist, left the company in November 2025 over a disagreement about architectural direction. He founded AMI Labs, a Paris-based startup, to pursue an alternative he calls the Joint Embedding Predictive Architecture, or JEPA. AMI Labs raised $1.03 billion in seed funding in March 2026 at a $3.5 billion pre-money valuation, the largest seed round in European startup history, betting specifically that world models, not larger language models, are the path to systems that can reason and plan reliably.
LeCun's argument is structurally different from Sutton's, even though both reject language-model-centrism. Sutton's objection is about learning dynamics: real intelligence requires an agent that keeps learning from lived experience, and a model trained once on a frozen corpus cannot do that no matter how large it gets. LeCun's objection is about what the model represents in the first place. In his framing, large language models are trained to predict the next token, which means they are trained to reconstruct surface form, and nothing in that objective requires the system to learn how physical events cause one another. JEPA is built to predict abstract representations of future states directly in a latent, embedding space, rather than reconstructing pixels or tokens, on the theory that causal, physical structure lives in that representation, not in the surface form a generative model is graded on. Two Turing Award winners, from two different subfields, converging independently on the same diagnosis, is a stronger signal than either one alone.
Semiotics did not wait for reinforcement learning, or for the internet, to notice that a symbol and the thing it stands for are not the same object. Ferdinand de Saussure's structural linguistics, developed in the early twentieth century, split the linguistic sign into a signifier, the sound or written form of a word, and a signified, the concept it points to, and argued the connection between the two is arbitrary, a matter of convention rather than necessity. There is nothing in the sound "tree" that resembles a tree. The word works because a community of speakers agreed it would, not because the sign contains the meaning inside itself. Charles Sanders Peirce, working independently in the same period, went further with a three-part model: a sign, an object it refers to, and an interpretant, the understanding a mind constructs when it encounters the sign. Peirce also distinguished icons, which resemble what they represent, indexes, which are causally connected to it, and symbols, which are purely conventional, of which most human language is the last kind.
The relevant point for this paper is narrow but load-bearing: in both traditions, the sign itself carries no meaning on its own. Meaning is something a system, historically a human mind embedded in a community and a world, constructs in response to the sign. A word is an address, not the thing at the address. This is precisely the distinction Sutton is gesturing at when he separates language from intelligence. Predicting the next plausible symbol in a sequence, which is what a language model is trained to do, is a claim about how signs are conventionally arranged relative to one another. It is not, by itself, a claim about the object or the interpretant those signs are supposed to be standing in for. A system can become extraordinarily good at arranging signifiers correctly, the way an unusually well-read parrot might, without that competence closing the gap semiotics identified between the signifier and the signified.
A large language model, mechanically, is a system trained to predict the next token in a sequence, conditioned on the tokens before it. Every token is a signifier in Saussure's sense: an arbitrary, conventionally assigned symbol. The model becomes extremely good at learning which signifiers tend to follow which others, across an enormous corpus of human-produced sign sequences, and that competence is genuinely impressive and, per Sutton, genuinely a scientific breakthrough. What the architecture does not do, because nothing in its training objective asks it to, is hold a persistent representation of the signified independent of the specific sentence being generated. The "meaning" of a token, in a transformer, is entirely relational: it is defined by its statistical relationship to other tokens in context, not by reference to a stable structure that exists whether or not that sentence is ever produced.
Essence's Meaning Coordinate system, 256 coordinates spanning four realms and thirty-two groups, is not an attempt to build a better guesser. It is built on a different premise entirely: that a human's intent can be declared as a position in a structured meaning space, by the human, and referenced by the machine, rather than reconstructed by the machine from whatever words happened to be used. This is a conceptual and architectural framing, not an empirical proof that the coordinate space captures more meaning than a token space does; that would be a different kind of claim than this paper is making. The claim here is narrower and, for what this paper is actually about, more relevant: a system whose primary unit is a declared coordinate does not need to interpret intent at all, because the interpretation step has already happened, by the person, at the moment they specified it. A system whose primary unit is the next-token probability has no such option. It has to infer the signified from the signifier every time, because nothing was ever declared to it directly. Intent-native computing generates machine instructions at runtime from that declared coordinate, rather than compiling code or predicting tokens from a corpus, which is the specific sense in which this architecture removes a step Sutton's critique assumes is unavoidable.
The 256 Meaning Coordinates are openly published, organized across four realms and thirty-two groups. A declared intent decomposes into a short sequence of them, the same way a sentence decomposes into words, except the sequence is the fact the machine executes against, not a guess it has to make. Seeing the decomposition happen against an actual sentence makes the distinction concrete faster than a paragraph can.
Try the live Meaning Coordinates demo →| Semiotic Layer | Token-Bound (LLM) Architecture | Meaning-Coordinate (PowerAptiv) Architecture |
|---|---|---|
| Primary unit of computation | The token, a signifier; probability distributions over what token follows | The coordinate, a position the human declared, not one the machine inferred |
| Relationship to the signified | Implicit, reconstructed statistically on each pass, never held as a standing fact | Explicit; the coordinate is the fact the instruction is generated against |
| Who resolves the interpretation | The model, probabilistically, on every call | The human, once, at the point of declaration; the machine only executes |
Section 04 described a single declared action resolving to a single instruction, and it would be a mistake to stop there, because it makes Essence sound smaller than it is. Sutton's Big World Hypothesis, the theoretical foundation Section 01 introduced, is not really an argument about tokens. It is an argument about world models: the world is too complex for any single agent to represent completely, so an agent's model of it is always a severe approximation, and the only way to keep that approximation useful is to keep updating it from lived experience. LeCun's JEPA program, introduced in Section 02, is aimed at a related but distinct target: not the learning dynamics, but the representation itself, predicting the causal structure of physical events in an abstract embedding space rather than reconstructing surface form. Both are, in their own vocabulary, saying a fluent token sequence is not a world model. It is a description that a mind, human, statistical, or otherwise, still has to turn into one.
Essence's answer to that problem is not "we don't need a world model because we don't guess." PowerAptivs include a category, Sapient, built specifically to run spacetime simulations tuned to the perspective of the entity experiencing them, adapting fidelity and scope as conditions change, and an Action, Emulator, that can govern an entire system simulation, spanning signal, sensor, spatial, and temporal categories at once, from a single declared intent. That is world simulation, not single-action execution, and PowerAptivs do not do it alone. The other seven Aptiv Types are what get simulated:
Together, those seven types are the declared, structured world; PowerAptivs are what simulates and executes against it.
That distinction matters more than a claim about representation, because it changes what continual learning, of either Sutton's or LeCun's kind, would even mean inside Essence. A BeliefAptiv's contents can change as new observations are declared into it, by a person or by a governed AI proposal that Synergy checks before it is recorded. A simulation running under the Sapient category can adapt its fidelity and scope as conditions change mid-run. Both of those are real changes happening over time, inside a world Essence is actively simulating. What it does not have, and is not claiming to have, is either Sutton's or LeCun's specific mechanism: an agent whose own internal representation, whether shaped by lived experience or by self-supervised prediction, shifts in a form that is not individually declared or individually attributable by design. Essence's world updates on declared, attributable, governed events. Sutton's and LeCun's update on training signal no one has to hand-declare, one from an agent's own experience, the other from prediction over large observational corpora. Those are three different answers to the same underlying problem, not three different amounts of progress toward one answer.
| World-Model Question | Sutton: Implicit, Learned From Experience | LeCun: Implicit, Learned From Prediction | Essence: Explicit, Declared |
|---|---|---|---|
| Where does the world model live | Inside the agent's parameters, shaped by accumulated experience | Inside a latent embedding space, shaped by self-supervised prediction over observations | Across governed Aptiv records: entities, events, scenes, beliefs, hypotheses |
| How the model updates | Through the agent's own experience, discovering abstractions no one declared to it | Through gradient-based training on video or sensor corpora, without hand-labeled targets | Through declared, attributable events: a person or a governed AI proposal Synergy checked |
| What failure looks like | A plausible-sounding output shaped by a world model no one can inspect | A confident prediction in embedding space that a probe reveals was never grounded in cause and effect | A declaration that was wrong or incomplete, traceable to who made it and when |
Detection ≠ Determination has always described a division of labor: GenAI proposes, Synergy governs. Read through semiotics, that division of labor is not a claim about which system is a better interpreter. It is a claim about which layer is allowed to guess at all. A Detection-layer posture trusts the sign itself, treating a plausible-sounding token sequence, or a confident embedding-space prediction, as sufficient evidence of correct output. That is exactly the posture Sutton and LeCun are each pointing at, from different disciplines, when they separate fluent output from real understanding of the world, and it is also, separately, not the posture Essence was built around. A Determination-layer posture does not ask GenAI to interpret intent more intelligently. It asks whether the proposed action matches a coordinate the human already declared, checked against a governed record, before anything executes. GenAI's output is treated as a proposal, one candidate reading of a signifier, not as the meaning itself. The system does not need to get better at guessing what that proposal meant, because the record it is checked against already holds what was actually meant.
Sutton is right that language models should not be mistaken for the whole of intelligence, and semiotics gives that observation a discipline older than either reinforcement learning or the transformer architecture: a sign is not the thing it signifies, and getting extremely good at arranging signs is not the same accomplishment as constructing the meaning those signs were meant to carry. Where this paper departs from the current conversation entirely is that it does not read that observation as a call to build a smarter interpreter. Essence's answer to the gap Sutton names is not a claim to have closed it with a better architecture for guessing. It is a decision not to guess. Humans declare intent directly, structurally, into a governed coordinate space. The machine's job is to execute against that declaration and to govern whether it was authorized, not to reconstruct what the human probably wanted from a sequence of words.
That is a narrower claim than solving intelligence in Sutton's terms, and it is a different claim, not a smaller version of the same one. This paper does not answer Sutton's demand for continual learning from experience, or LeCun's demand for a self-supervised world model, and it is not trying to. It does not claim a fixed set of 256 Meaning Coordinates exhausts the space of human intent, any more than a finite vocabulary exhausts the space of things that can be said with it. What it claims, alongside Section 05's fuller picture, is that a declared coordinate checked by Synergy before it acts is the same mechanism whether the action is a single instruction or an update to a simulated world: the machine is never asked to guess what was meant. It only has to check what was declared.
Language is a real and remarkable capability, and treating a breakthrough in it as the whole of intelligence is, as Sutton argues, a category error worth correcting. LeCun makes a related argument from representation rather than learning dynamics: a system trained to reconstruct surface form is never required to learn the causal structure underneath it. Semiotics gave both objections a shared name a century before either field existed: a sign is not its meaning, and no amount of skill at arranging signs, or predicting them, closes that gap on its own. The deeper target both critiques share is the requirement for a real world model, and this paper's answer to it is not to sidestep the problem. PowerAptivs simulate, using a spacetime-aware category built for exactly that; StoryAptivs, ExperienceAptivs, BeliefAptivs, MindAptivs, RecordAptivs, IdeaAptivs, and SignalAptivs are the declared, structured world that simulation runs against. That world updates, the same way Sutton and LeCun both say a world model must, through observations, predictions, and changing conditions. It updates on declared, attributable, governed events instead of on experience an agent had, or a pattern a network absorbed, that no one had to individually account for. Neither critique is a challenge this paper is trying to answer point for point. Both are confirmation that a world model was always going to be necessary, and evidence that the model does not have to be a black box, statistical or latent, to be real.
Request Platform Access → Full White Paper Series