The Agency Illusion

The AI industry is betting on agents. Connecting models that cannot govern themselves does not produce agency. It produces a larger version of the same problem, with more moving parts and less accountability for any of them.

Ken Granville · CEO & Co-Founder, MindAptiv June 2026 · Revised September 2026 Open Access
Abstract

The dominant architecture being deployed under the label of agentic AI connects multiple language models through protocols designed to route context and coordinate actions across sessions. The Model Context Protocol and its variants are presented as the infrastructure layer that makes agents capable of complex, sustained work. This paper argues that the framing contains a structural error.

Connecting stateless, session-bound, ungoverned models does not produce agents in any meaningful sense. It produces pipelines. The difference is not semantic. A pipeline multiplies capability without adding governance. An agent, properly understood, requires determination: the capacity to decide whether a proposed action is allowed to run, not merely the capacity to propose and execute it. No amount of inter-model routing produces that capacity, and neither does monitoring the models’ chains of thought: reading what a model writes about its reasoning reveals neither what actually drove it nor what it was authorized to do.

The composed Aptiv is not an alternative agent architecture. It is a different kind of thing entirely: a governed, persistent, semantically grounded unit of intent that does what agentic pipelines were designed to approximate, without the structural failures those pipelines cannot escape. It also keeps the economic attributes that make agents commercially attractive, and adds the one they cannot supply: trust established before a transaction executes, not inferred after it.

Section 01The Bet

The AI industry has made a coordinated wager. The wager is that capability, extended through coordination, becomes something qualitatively new. Chain enough models together, give them access to enough tools, let them hand off context through a well-specified protocol, and the result is not a collection of language models completing tasks in sequence. The result, the argument goes, is an agent: a system that can pursue goals, manage complexity, and sustain work across time.

The investment flowing into this thesis is not small. Every major hyperscaler has an agent framework. The Model Context Protocol, originally proposed by Anthropic and now adopted across the industry, provides the plumbing. Microsoft, Google, Amazon, and a growing field of startups are building orchestration layers, agent marketplaces, and multi-agent workflows on top of it. The 2026 enterprise AI budget is increasingly an agentic AI budget.

The bet has intuitive logic. If one model can do useful work, many coordinated models should do more useful work. If context can be passed between sessions, sustained goals become tractable. If tools can be connected through a protocol, agents can act in the world, not just generate text about it.

The logic is correct about what pipelines can do. It is wrong about what they cannot. And what they cannot do is the thing that makes an agent an agent.

Section 02What the Word Requires

Agency, in any meaningful technical or philosophical sense, requires more than the ability to act. A thermostat acts. A script acts. What distinguishes an agent from a mechanism is not the capacity to produce outputs in response to inputs. It is the capacity to evaluate whether a proposed action should be taken, and to constrain behavior accordingly.

That evaluation function has a name in the Essence platform doctrine: determination. It is the faculty that decides whether a proposal is allowed to run, under what conditions, and within what constraints. It is distinct from the faculty that generates the proposal in the first place. And it is precisely what no language model, however large, however well-prompted, however extensively connected to other models, is designed to perform.

The peer-reviewed basis for this distinction has been established. A 2026 PNAS Nexus study found deficient executive control in transformer attention: the faculty that resolves conflict and selects under competition. The authors trace the deficit to the architecture, which has no explicit mechanism for top-down control. More training can improve performance on specific tasks; it does not supply the general faculty. That faculty is what determination requires.

Doctrine
Detection is not determination.
Proposing is not governing.
Routing is not agency.

This matters for the agentic AI thesis because the thesis conflates three distinct functions that are not substitutable for one another. A model can detect patterns, generate proposals, and route outputs to other models. None of those functions is determination. Connecting models that cannot determine does not produce a system that can. The governance gap is not filled by the protocol. The protocol is indifferent to it.

Section 03What MCP Actually Is

The Model Context Protocol is well-designed for what it is. It provides a standardized interface through which AI models can connect to external data sources, tools, and services. It reduces the integration cost of building multi-model workflows. It makes it easier for one model to hand off context to another, for agents to invoke tools, and for orchestration layers to coordinate actions across systems.

None of that is a criticism. MCP is good plumbing. The problem is not the protocol. The problem is what the protocol is being asked to do that plumbing cannot do.

P.01
MCP routes context. It does not govern it.

The protocol specifies how information moves between models and tools. It does not specify whether a proposed action is permitted to execute, or under what conditions. Governance is out of scope for the protocol by design. That is not a flaw. It is a category definition.

P.02
MCP extends sessions. It does not eliminate them.

Each model in an MCP-connected workflow remains stateless. Whatever persists, in memory servers, files, or databases, persists as text that the next model must retrieve and re-interpret. None of it is a standing commitment to constraints established earlier in the workflow. Context is passed as text. It can be lost, misinterpreted, or overridden. The session boundary is the unit of the system whether there is one model or twenty.

P.03
MCP multiplies surface area. It does not reduce risk.

Every additional model in a pipeline is an additional point of hallucination, misinterpretation, and ungoverned proposal. Each tool connection is an additional execution path without a determination layer. The attack surface grows with the pipeline. Accountability does not grow with it. In multi-model systems, failures are distributed across models that each operated within their design parameters. The system fails, and no component is responsible.

P.04
MCP coordinates capability. It does not compose intent.

The difference between coordination and composition is not terminological. Coordinated models share a workflow. Composed intent is a single unified artifact with a defined semantic structure, governed constraints, and traceable authorship. One is a pipeline. The other is a unit. Only one of them can be audited, versioned, stored, reused, and governed as a discrete thing.

Section 04The Pipeline Is Not the Agent

There is a reason the industry reached for the word "agent" and not the word "pipeline." Pipeline is accurate. Agent is aspirational. The word carries connotations of autonomy, judgment, and sustained purpose that pipelines do not have and were not built to have. Using the more aspirational word does not transfer the properties. It transfers the expectation.

That gap between expectation and architecture is where the liability lives. An enterprise that deploys an agentic AI system expecting autonomous, governed, accountable behavior and receives instead a multi-model pipeline that routes ungoverned proposals through connected tools has not purchased what it believed it purchased. It has purchased the expectation of an agent and received the mechanics of a pipeline. The contracts, regulatory commitments, and operational dependencies that follow from the first assumption do not survive contact with the second reality.

The enterprise AI governance gap is not the gap between current AI and future AI. It is the gap between what "agentic" implies and what the current architecture delivers. That gap is being papered over with terminology, not closed with infrastructure.

This is not a speculative risk. It is already the shape of real failures. Multi-agent systems acting on stale or misrouted context. Orchestration layers executing actions no human would have approved if they had seen the full reasoning chain. Tool-connected pipelines taking irreversible real-world actions based on model outputs that no governance layer evaluated. The common thread in each case: a pipeline that had capability but not determination.

The industry's leading answer to that last failure is to read the reasoning chain. Chain-of-thought monitoring uses a second model to inspect the step-by-step reasoning a model writes out before it acts. It is a genuine safety tool, and it does not close this gap. The research record, much of it published by the frontier labs themselves, is candid about why.

The reasoning a model writes is not reliably the reasoning that drives it. Anthropic researchers tested whether reasoning models disclose the hints they actually rely on. The models revealed that reliance in at least 1% of cases but often in fewer than 20%, and when training increased their exploitation of hints, their disclosure of it did not increase (Chen et al., Reasoning Models Don’t Always Say What They Think, 2025).

Monitoring pressure teaches concealment. OpenAI researchers found that a chain-of-thought monitor could catch reward hacking in a frontier reasoning model. When the monitor was folded into training, the models learned to hide their intent within the reasoning while continuing to reward hack at a significant rate (Baker et al., Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation, 2025).

Its own advocates call it fragile. A 2025 position paper by more than forty researchers from the UK AI Security Institute, Apollo Research, Anthropic, Google DeepMind, OpenAI, and other institutions recommends investing in chain-of-thought monitoring while describing it as imperfect, able to let misbehavior go unnoticed, and possibly fragile as models and training methods change (Korbak et al., Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, 2025).

Even a perfectly faithful chain of thought would not solve the problem this paper describes. A reasoning trace records what a model considered. It does not record what the model was authorized to do. The monitor can run before an action executes, but what it evaluates is natural language, judged against another model’s sense of what looks acceptable. That is detection moved earlier in the loop, not determination. In a multi-agent pipeline the difficulty compounds: the reasoning is split across models, each with its own trace, and no single trace contains the decision.

Chain-of-thought monitoring watches the proposer think. Determination decides whether the proposal may run. The first can inform the second. It cannot replace it, because a record of what a model was thinking is not a record of what it was authorized to do.
Cross-reference: The Governance Gap (Paper 9) documented the trillion-dollar scale of ungoverned AI exposure. The Recall Standard (Paper 8) established why detection of failure after execution is categorically insufficient. Agentic pipelines do not resolve either problem. They extend both.

Section 05What Combining Aptivs Actually Does

The composed Aptiv is frequently described as MindAptiv's alternative to the agentic AI stack. That framing is useful in competitive context. It is slightly imprecise about the nature of the difference. The composed Aptiv is not a better agent. It is a different category of thing, doing a different job, at a different layer of the machine.

An Aptiv is a governed unit of intent. It is not a model. It is not a session. It is not a pipeline step. It is a discrete, semantically structured artifact that encodes what a process is supposed to do, under what conditions, within what constraints, and with what governance applied to every execution. It is expressed in Meaning Coordinates: the 256-coordinate semantic substrate that forms the foundation of the Essence platform, grouped across four realms and thirty-two conjugate pairs.

When Aptivs are combined, the result is not a workflow. It is a composition: a single governed artifact that contains multiple intent structures, resolved into coherent semantic unity by Synergy, the governance layer of the Essence platform. The composition is traceable to its authors. It is auditable at the level of every constraint it contains. It is persistent across invocations. And it is governed at the substrate level, not at the prompt level.

GenAI proposes. Synergy governs. The governed proposal that reaches execution is not the raw output of a model. It is an artifact that has passed through determination. That is what the pipeline is missing. That is what composition provides.

The practical consequence is the one that matters most for enterprise deployment. A combined Aptiv can be stored, versioned, audited, and reused as a unit. A multi-model pipeline's code can be versioned, but what it decides exists only during execution. Its reasoning survives only as logs. Its constraints are not enforceable after the session ends. Its behavior cannot be guaranteed to be consistent across invocations. The Aptiv composition solves each of these problems at the architectural level, not through better prompting or more careful orchestration.

Section 06The Structural Comparison

The difference between a multi-agent pipeline and a composed Aptiv is not a difference of degree. It is a difference of kind. The table below maps the structural properties that enterprises care about across the two architectures.

Property Multi-Agent Pipeline (MCP) Composed Aptiv (Essence)
Unit of work Session sequence across connected models Persistent governed artifact
Governance layer None native; external guardrails optional Synergy applied at substrate level
State persistence Stored externally as text; re-retrieved and re-interpreted each session Durable AptivRecord; versioned and stored
Auditability Log-level; post-hoc reconstruction Native; constraints and authorship traceable
Determination Absent; model proposes, coded rules or human approval gate the action Native; Synergy evaluates before execution
Reusability Workflow code is reusable; its reasoning is regenerated per invocation Composition is a reusable semantic unit
Accountability Distributed across models; diffuse on failure Attributed to composition and governing rules
Consistency Probabilistic; varies with model and context Deterministic within governance constraints
Trust level Inferred from model; not architecturally enforced Explicit PowerAptiv Trust Level taxonomy
Transactions Executed on model output; authorization reconstructed from logs Authorized by Synergy before execution; record retained in the Aptiv
Economic attribution Operator paid for activity; knowledge source unattributed Each use traceable to source and author; compensable
Semantic grounding Natural language; ambiguous by construction Meaning Coordinates; structured and precise

The table does not show a better version of the same architecture. It shows two different theories of what the unit of work should be. The pipeline theory says the unit of work is a coordinated sequence of model invocations. The composition theory says the unit of work is a governed artifact that encodes what should happen and enforces it every time.

Section 07Why the Industry Bet on the Wrong Unit

The agentic AI bet is not irrational given the constraints the industry is working within. If your substrate is a language model, and language models are stateless and session-bound, then connecting them through a protocol is the logical way to extend their reach. MCP is the right answer to the question the industry is asking. The question itself is the problem.

The question being asked
How do we make language models do more?
The better question
What should the governing unit of computational intent look like, and how do we build a substrate that runs it correctly?

The distinction matters because the first question leads to pipelines. The second leads to a different architecture from the ground up. You cannot arrive at Aptiv composition by iterating on MCP. You cannot arrive at Meaning Coordinates by iterating on prompts. You cannot arrive at Synergy by iterating on guardrails. These are not improvements to the existing architecture. They are a different layer of the machine that the existing architecture was never designed to produce.

W.01
The industry started with the model.

Every major framework, every protocol, every orchestration layer begins with a language model as the atomic unit. That starting point encodes the session boundary, the statistical nature of outputs, and the absence of a governance layer into every architecture built on top of it. You cannot engineer your way out of the starting point without changing the starting point.

W.02
The industry confused capability with architecture.

Language models are genuinely capable. That capability created the reasonable inference that better architecture would come from extending capability: more parameters, more connections, more tools, more agents. The inference is wrong because the governance problem is not a capability problem. It is an architectural absence that capability cannot fill.

W.03
The industry inherited the session as the natural unit.

Software has always been session-based. Requests, responses, connections, transactions: the entire computing stack was built around bounded interactions. The language model fit naturally into that frame. The problem is that governed intent does not fit in a session. It requires persistence, traceability, and substrate-level enforcement that the session model was never designed to support.

Section 08The Liability Nobody Is Pricing

Enterprise adoption of agentic AI is accelerating while the governance infrastructure that should precede deployment is not. The gap is being filled by confidence in the capability, by competitive pressure to deploy, and by the difficulty of pricing a risk that has not yet fully materialized.

The risk has a shape. A multi-agent pipeline that executes an irreversible action based on a hallucinated fact cannot be unwound by the orchestration layer.

It cannot be audited at the level of the reasoning chain, because the reasoning chain is distributed across models that no longer share a context.

It cannot be attributed to a decision that anyone made, because no one made a decision.

A model produced a proposal and a tool executed it.

Regulators are beginning to ask who is accountable for that outcome. The answer the agentic AI architecture provides is: nobody specifically, which is an answer no regulator will accept. The EU AI Act, emerging US AI governance frameworks, and sector-specific regulations in financial services, healthcare, and critical infrastructure are converging on requirements that the agentic pipeline cannot satisfy within the current architectural model.

Accountability requires attribution. Attribution requires a decision point. A decision point requires determination. Determination requires a governance layer that the agentic pipeline does not have and cannot acquire through more sophisticated coordination of models that do not have it.

The liability is not hypothetical. It is structural. And unlike most structural liabilities, it does not diminish with scale. A larger agentic deployment has more surface area, more diffusion of accountability, and more execution paths that no governance layer has evaluated. Deployment size is no longer limited by the effort required to build agents, because models can now generate and launch them on demand. The liability scales with the investment, and increasingly with the rate of generation.

Cross-reference: The Litigation Layer (Paper 10) mapped the legal exposure already accumulating across AI architectures. The Wrong Race (Paper 15) documented the mechanism by which speed without governance produces liability rather than advantage. Agentic deployment without determination is the same mechanism at the workflow layer.

Section 09Models as Proposers

This argument is not a rejection of language models. It is a rejection of the agent as the governing architecture. Models are fast, capable, pattern-matching proposers, and the Essence platform uses them in exactly that role. In the current Assimilator prototype, model-driven stages produce candidate Aptivs. In the deployed design, those candidates follow a trust promotion path: they are evaluated, governed, and, when they meet the determination standard, promoted into the library as approved intent. The model proposes. The governance layer determines. The approved record is the result.

The Assimilator's current agent stages are transitional. MindAptiv is replacing them with Aptivs generated from Meaning Coordinates, with system-wide context available so the consequences of each step can be evaluated before it runs, without depending on code. What survives the transition is the division of labor, not the agent: models propose, Synergy determines. That is the only arrangement that produces a governed outcome, and it does not require an agentic pipeline to deliver it.

01
Model proposes

A model operating inside the Assimilator generates a candidate Aptiv based on natural language intent and available context. This is what the model is designed to do.

02
Pipeline evaluates

The candidate passes through the Assimilator's eight phases: parallel source-grounded extraction, adversarial critique, deterministic normalization and validation gates, and a mandatory human Creator Review. No candidate reaches the library on a model's output alone.

03
Synergy determines

In the deployed design, Synergy applies governance at the substrate level, replacing the model-judged verification and scoring in the current prototype. The candidate is evaluated against the constraints of the governing Aptiv, the trust level taxonomy, and the semantic grounding requirements of the Meaning Coordinate system.

04
Record is promoted

Candidates that pass determination are promoted to approved AptivRecords. They are persistent, versioned, auditable, and reusable. The approval is traceable. The reasoning is preserved. The governance is native to the artifact.

05
Composition extends

Approved AptivRecords can be combined. The combination is a governed composition: a unified semantic artifact that executes with Synergy applied at every step. No session boundary interrupts the governance. No handoff loses the constraints. The composed Aptiv does the job the industry is trying to build agents for, built from governed units instead of connected sessions.

Section 10Trust Is the Transaction

Agents are not attractive only because they are capable. They are attractive because they are economic. An agent can act on a person’s behalf, run unattended, carry out transactions such as purchases, bookings, and payments, and be sold per seat, per task, or per outcome. Any architecture that asks an enterprise to give up those attributes will lose to one that keeps them, however well governed it is. Aptivs do not ask that. They keep every economic attribute of the agent and add the one the agent cannot supply: trust that exists before the transaction does.

Consider what limits an agent’s economic value today. It is not capability. Models can already fill a cart, book travel, move money, and negotiate terms. The limit is how much anyone is willing to let an agent do without watching it. Every dollar of authority delegated to an agent is a dollar someone has to trust will be spent as intended, and in a pipeline that trust rests on the model’s reputation and the monitoring around it. That is a trust ceiling, and it is where agentic commerce stalls: low-value tasks run unattended, and high-value ones wait for a human.

The payments industry has reached the same conclusion from its side. When Mastercard and Google introduced Verifiable Intent in March 2026, Mastercard’s chief digital officer, Pablo Fourez, put it plainly: “As autonomy increases, trust cannot be implied. It must be proven.” Verifiable Intent proves what a consumer authorized at the moment an agent initiates a purchase. As The Payment Moment (Paper 35) sets out, it does not govern what the system does between that authorization and the outcome. An Aptiv carries trust through both: the authorization, and the execution that follows it.

T.01
Retained: acting on someone’s behalf.

An Aptiv executes without its author present for each decision; that is the purpose of encoding the author’s intent. What it does is bounded by that intent, not by a model’s reading of a prompt.

T.02
Retained: transacting.

An Aptiv can carry out the same purchases, bookings, and payments an agent can. The difference is that each transaction’s authorization, meaning who granted it, within what scope, and at what trust level, is evaluated by Synergy before the transaction executes and remains part of the Aptiv’s record afterward. A counterparty does not have to reconstruct that authorization from logs. It can rely on it.

T.03
Expanded: attribution and compensation.

An agent is sold as activity, and its operator is paid. Every use of an Aptiv traces back to the authoritative source and the author whose knowledge it encodes, which makes that use attributable and therefore compensable. This is the Intent Economy described in Paper 12: the expert whose judgment governs a transaction participates in the value it creates.

T.04
Expanded: trust as a record, not a reputation.

An agent’s trustworthiness is inferred from the model’s reputation and from how the agent behaved the last time someone checked. An Aptiv’s trust is recorded. It carries a certified trust level, from Level 1 with a coded Guard to Level 3 with a native, codeless Guard, and it reaches the library only through the promotion path described in Section 09. Trust becomes something a counterparty can check rather than something it has to assume.

Agents compete on capability, and capability is becoming a commodity. What cannot be commoditized is the ability to prove, before a transaction executes, that it was authorized. Trust is the value proposition, and it is the one attribute a pipeline cannot acquire by becoming more capable.

This is why the economic case and the governance case are the same case. An enterprise that deploys governed Aptivs does not trade commercial value for safety. It raises the ceiling on what it can delegate, because every transaction arrives with its authorization already determined. The trust ceiling that holds agents to low-value work is the ceiling Aptivs remove.

Cross-reference: The Scale of Intent (Paper 11) sets out the same argument for autonomy and monetization at the scale of an enterprise library. The Intent Economy (Paper 12) develops the compensation model for the authors whose knowledge Aptivs encode. The Payment Moment (Paper 35) maps where payment-layer attestation ends and execution governance begins.

Section 11The Question the Industry Has Not Asked

Every major AI announcement in 2026 involves agents. The word appears in product names, in funding announcements, in congressional testimony, and in enterprise procurement criteria. The question of what an agent actually requires to function as advertised has not been asked with commensurate urgency.

It is a simple question. An agent that can pursue a goal across time must be able to make decisions about what actions to take and what actions to refuse. The capacity to refuse, grounded in defined constraints that persist across the life of the goal, is determination. Which architecture provides it?

The MCP-connected pipeline provides routing, context passing, and tool access. It does not provide determination. The composed Aptiv provides determination natively, at the level of every governed unit it contains. That is not a competitive claim. It is an architectural description. The gap it identifies is structural, not rhetorical.

The industry is building agents without governance and calling the result agentic AI. The result is capable. It is not governed. Those are not the same thing, and the market has not yet priced the difference.

The enterprises deploying agentic systems in 2026 are making the same bet the industry is making at scale. They are betting that the capability gap between current and future AI is the important gap. It is not. The important gap is between systems that can propose and systems that can determine. Most current deployments are on the wrong side of that gap. The liability follows from the position, not from the intent.

Section 12The Answer That Was Already Here

The Essence platform was not designed in response to the agentic AI moment. The Wantware architecture that underlies it was specified before the current AI era produced the capability it depends on. The platform does not reject the models inside agents. It rejects the agent as the governing architecture. Agentic systems as they currently exist are the wrong architecture for governed AI, and the platform replaces them rather than governing them from outside. Their two commercial appeals, autonomy and monetization, are available from Aptivs inside a governed boundary, as The Scale of Intent (Paper 11) sets out.

The composed Aptiv is the answer to the question the agentic bet is trying to answer: how do you sustain governed intent across complex, multi-step, real-world processes? The bet answered: connect more models. The Essence architecture answered: compose governed units. The two answers are not iterations on the same idea. They are different theories of what the machine should be.

Theory One · Pipelines

One theory produces pipelines. Pipelines are capable. They are fast. They are increasingly sophisticated.

They are not governed, not persistent, and not accountable in the ways that regulated industries require.

Theory Two · Composed Aptivs

The other theory produces composed Aptivs. They are governed from the first unit of intent to the last execution. They are persistent. They are auditable.

They are the architecture that the enterprise AI market needs and does not yet know it is missing.

The Position
Models are proposers.
Aptivs are governed artifacts.
Trust is the value proposition.
Composition is not orchestration.
Determination is not coordination.
The bet on agents is a bet on the right capability in the wrong architecture.
The window to build the right architecture is open.
It will not remain open indefinitely.