The Context Fatigue Ceiling

Why Frontier Agents Fail on Long-Horizon Work, and Why Scaling Cannot Fix It

Ken Granville · CEO & Co-Founder, MindAptiv July 2026 Open Access
Abstract

Carnegie Mellon University's Gym-Anything paper, published April 2026,1 is the most rigorous large-scale measurement of computer-use agents on realistic enterprise software to date. Its CUA-World benchmark spans 10,215 tasks across 200 applications covering all 22 major U.S. occupation groups.

The headline result is unambiguous: on long-horizon tasks, every frontier model tested achieved a pass rate below 8% under standard cost constraints. The paper identifies the primary failure mode and gives it a name: context fatigue. After processing hundreds of thousands of tokens, agents lose track of what they were supposed to accomplish. They declare tasks complete prematurely, substitute placeholder data for real outputs, and abandon their original intent mid-trajectory.

The authors treat this as a technical problem to be solved with better training, longer context windows, and test-time auditing. This paper argues the diagnosis is correct and the prescription is incomplete.

Context fatigue is not a memory problem. It is an architecture problem. An agent that holds intent inside its context window will always be vulnerable to losing it. A platform that holds intent outside the model, as a governed, persistent representation evaluated before execution, cannot forget what the work is for. The fix is not a longer window. It is a different substrate.

Disclosure
The Gym-Anything paper discloses that the agents used to construct the CUA-World benchmark environments were Claude Opus 4.5 and Claude Opus 4.6, run via Claude Code. Claude Sonnet 4.6 is also one of the models evaluated in the benchmark results cited here. MindAptiv builds on Anthropic's API infrastructure. These relationships are stated for transparency; they do not alter the analysis that follows.

Section 01What the Benchmark Actually Proved

The hard problem in AI agent research has always been measurement. Existing benchmarks test narrow, short tasks: click a button, fill a form, navigate a website, configure a system setting. These are useful calibration tools for individual capabilities. They are a poor proxy for whether an agent can do real work.

The CMU team's response to this problem is methodologically significant.5 Rather than handcrafting a fixed set of tasks, they used agents to build the benchmark itself. A coding agent writes setup scripts, installs software, loads real-world data, and configures each application to a realistic starting state. An independent audit agent then verifies the setup against a quality checklist, issuing corrections before any task is attempted. The result is CUA-World: 10,215 tasks across 200 applications, grounded in a taxonomy derived from U.S. GDP data by occupation group, spanning domains from medical imaging analysis to enterprise resource planning to legal document management.1

This is the first benchmark that approximates the shape of actual professional work. Not a simplified proxy. Not a curated toy environment. Software that real workers use, populated with realistic data, requiring multi-step workflows that resemble what the work actually demands.

The results from that benchmark are reproduced in the table below, drawn directly from the paper. All figures cited here should be verified against the primary source at arxiv.org/abs/2604.06126.1

Model Type Benchmark Steps / Cap Pass Rate
CUA-WORLD-TEST (Standard Benchmark)
Gemini 3 Flash Proprietary CUA-World-Test 200 steps 22.6%
Kimi-K 2.5 Open-source CUA-World-Test 200 steps 12.8%
Qwen3-VL-4B (base) Open-weight CUA-World-Test 200 steps 3.9%
Qwen3-VL-2B (distilled) Open-weight CUA-World-Test 200 steps 4.4%
Qwen3-VL-2B (base) Open-weight CUA-World-Test 200 steps 1.6%
CUA-WORLD-LONG (Long-Horizon Benchmark)
Gemini 3 Flash Proprietary CUA-World-Long 500 steps / $5 7.5%
Kimi-K 2.5 Open-source CUA-World-Long 500 steps / $5 5.5%
Claude Sonnet 4.6 Proprietary CUA-World-Long 500 steps / $5 6.0%
GPT-5.4 Proprietary CUA-World-Long 500 steps / $5 3.0%
GPT-5.4 Proprietary CUA-World-Long 2,000 steps / uncapped 27.5%
Gemini 3 Flash Proprietary CUA-World-Long 2,000 steps / uncapped 11.5%
Gemini 3 Flash + TTA Proprietary CUA-World-Long 2,000 steps / uncapped 14.0%
Source: Aggarwal, Neubig, Welleck. "Gym-Anything: Turn any Software into an Agent Environment." arXiv:2604.06126, April 7, 2026. Tables 3 and 4. TTA = Test-Time Auditing. Claude Sonnet 4.6 and GPT-5.4 do not appear in Table 3 (CUA-World-Test); the paper evaluated them only on CUA-World-Long. GPT-5.4 exhausts its $5 budget in fewer than 100 steps on Long tasks, explaining its 3.0% capped result vs. 27.5% uncapped. Kimi-K 2.5 is described as open-source in the paper. Qwen3-VL models are open-weight. Model names are as cited in the paper; verify all figures against the primary source.

Two numbers in this table require additional attention because they reframe the others. GPT-5.4's capped pass rate of 3.0% is not primarily a capability failure. The paper notes the model exhausts its $5 cost budget in roughly 150 steps, well below the 500-step ceiling. Its uncapped result of 27.5% at 2,000 steps is the strongest in the study. The cost cap is a real-world constraint, not an artifact of the benchmark, and it reveals something important: the most capable model is also the most expensive to run on long tasks. Neither raw capability nor cost efficiency is sufficient.

The second number that reframes the table is the Test-Time Auditing result. Adding an external audit agent to Gemini 3 Flash improves its uncapped long-horizon pass rate from 11.5% to 14.0%. The paper presents this as a practical improvement. It is also an inadvertent structural finding, and Section 03 takes it up directly.

The most rigorous benchmark of realistic enterprise software work yet published shows every frontier model solving fewer than 8% of long-horizon tasks under standard cost constraints. This is not an edge-case result from a toy environment. It is a measurement of the gap between agent capability and real work.

Section 02The Failure Mode Has a Name

The Gym-Anything paper does not simply report the pass rates. It diagnoses the failure. The authors examined agent trajectories on failed long-horizon tasks and identified a consistent pattern they call context fatigue, citing earlier research on the phenomenon under the label "Lost in the Middle."2

The description is precise. After processing hundreds of thousands of tokens across a long trajectory, the agent loses track of what it was originally trying to accomplish. The degradation is not random. It takes characteristic forms.

The agent declares the task complete while the application is at the wrong screen. It substitutes placeholder or sample data for the real outputs the task required. It completes a subset of the required steps and stops, without recognizing that the remaining steps were part of the same task. In each case, the original intent is present somewhere in the context window. The agent simply can no longer act on it reliably.

The paper speculates on the mechanism, consistent with the "Lost in the Middle" literature: information in the middle of a very long context is retrieved less reliably than information near the beginning or end. As a trajectory grows, the original task specification migrates toward the middle, and the agent's ability to act on it degrades. The longer the task, the more severe the degradation.

Context fatigue is not forgetting. The original intent is still in the window. The agent simply cannot act on it reliably once the window is long enough. That is a different problem than memory, and it has a different fix.

This distinction matters because the natural response to context fatigue is to make the context window longer. If the agent loses intent because the original task specification slides toward the middle of a long window, extend the window so the specification stays closer to the boundary. This is the prescription the scaling paradigm reaches for instinctively. It is also, as Section 04 argues, the wrong prescription, because it treats a structural problem as a capacity problem.

Cross-reference: Paper 13: "The Session Illusion" | Paper 17: "The Agency Illusion"

Section 03The Auditing Finding Is a Structural Admission

The most consequential result in the Gym-Anything paper is not the pass rates. It is the Test-Time Auditing experiment, and what it reveals about the architecture the authors were not trying to describe.

The setup is straightforward. After a primary agent completes its trajectory, a second model reviews the trajectory against the original task specification, identifies what was missed or done incorrectly, and feeds that assessment back. Under the 2,000-step uncapped condition, Gemini 3 Flash improves from 11.5% to 14.0% pass rate with this addition. The authors present this as a practical technique for improving agent performance at test time.

But notice what the technique actually does. It externalizes the memory of the original intent. The primary agent, having processed a long trajectory, has lost reliable access to what the task required. The audit agent, reading the trajectory fresh against the original specification, re-introduces the intent from outside the primary agent's degraded context. The improvement in pass rate is a direct measurement of how much intent the primary agent had lost by the end of its trajectory.

This is the same structural move Paper 213 identified in the AI governance debate. Just as the entity-governance proposal reaches for continuous, AI-facilitated auditing because periodic human audits are insufficient, the Gym-Anything paper reaches for an external audit agent because the primary agent's internal state is insufficient. In both cases, the prescription is a governance layer added after the fact. In both cases, the underlying diagnosis points at the same missing architecture: a substrate that holds intent persistently, outside the executing model, from the beginning of the task rather than as a correction at the end.

Test-Time Auditing improves pass rates by re-introducing intent the primary agent had lost. That improvement is a measurement of the context fatigue deficit. It is also a proof by necessity: the system needs a layer that holds intent across the trajectory. The question is whether that layer arrives as a post-hoc correction or as a governing substrate from the start.
Cross-reference: Paper 21: "The Missing Substrate" | Paper 16: "The Ledger That Is Intent-Driven"

Section 04Why Scaling Cannot Fix This

The scaling response to context fatigue is not irrational. If the agent loses intent because the task specification is buried in a long context, a longer context window keeps the specification closer to retrievable positions. A better-trained model retrieves middle-context information more reliably. Both improvements are real and measurable. The uncapped GPT-5.4 result of 27.5% is evidence that removing cost and step constraints substantially improves performance. Progress on context length and retrieval fidelity will continue.

But the progress has a ceiling, and the ceiling is structural. Here is why.

A model holding intent inside its context window must carry that intent through every step of the trajectory. Every tool call, every screenshot observation, every intermediate result adds tokens to the window. The task specification, fixed at the beginning, does not grow; everything else does. The ratio of task-specification tokens to total context tokens decreases monotonically as the trajectory extends. No context length eliminates this dynamic; it only shifts the point at which degradation becomes severe. The "Lost in the Middle" problem2 is not a bug to be patched. It is a property of the architecture.

There is a second, independent ceiling: cost. The uncapped results in the Gym-Anything paper are not deployment conditions. They are laboratory measurements. At 2,000 steps per task, the per-task inference cost for frontier models is not commercially viable for the class of work CUA-World tests. The gap between the capped pass rates and the uncapped pass rates is not a benchmark artifact. It is a preview of the real constraint that enterprise deployment will impose. More capable models are more expensive per step. Long-horizon tasks require more steps. Both trends move in the direction of higher cost, not lower.

Intent-native architecture addresses both ceilings simultaneously, because it moves intent out of the context window entirely. In the Essence® platform, the wantverse holds what the user or system wants to accomplish as a governed, persistent representation that exists independently of any individual model invocation. Synergy® evaluates each proposed action against that representation before execution. The model does not need to remember the original task specification in its context window, because the task specification is not in the context window. It is in the substrate.

A 2,000-step trajectory does not erode the governing intent, because the governing intent was never inside the trajectory to begin with.

01
Intent is declared once, to the platform

The user or calling system expresses what is to be accomplished. This declaration is stored as a governed representation in the wantverse, not as tokens in a model's context window.

02
Each action is evaluated against the governing intent

Before any step executes, Synergy® evaluates the proposed action against the persistent intent record. The model proposes. The substrate determines whether the proposal is consistent with what the task actually requires.

03
The trajectory cannot drift from its purpose

Context fatigue is a failure of retrieval: the agent cannot reliably access intent buried in a long window. Intent held outside the window cannot be buried in it. The governing representation is not subject to the "Lost in the Middle" dynamic, because it is not in the middle of anything.

04
Every governed determination is recorded

Each evaluated action produces an AptivRecord: what was intended, what was permitted, what executed, and under what constraint. The trajectory produces a governed ledger, not a log. Audit, if required, verifies the ledger rather than reconstructing intent from a long context after the fact.

The Gym-Anything paper shows that Test-Time Auditing recovers some of the intent the primary agent lost. That recovery comes at the cost of an additional model invocation at the end of a trajectory that has already consumed thousands of steps. The intent-native alternative is not a correction applied at the end. It is a governing constraint applied at every step. The difference is not marginal. It is the difference between fixing drift after it has accumulated and making drift architecturally impossible.

A longer context window does not eliminate context fatigue. It moves the ceiling. Intent held outside the context window cannot fatigue, because it is not subject to the retrieval dynamics that produce fatigue. The fix is not a larger window. It is a governed substrate that does not rely on the window at all.

Section 05The Open-Source Argument Does Not Close the Gap

A predictable response to the cost ceiling documented above is to substitute smaller, open-weight models for frontier proprietary ones. If GPT-5.4 costs approximately $18 per long-horizon trajectory uncapped, and Gemini 3 Flash costs approximately $16, the argument runs: use a fine-tuned open-weight model instead, run it locally, and eliminate the per-step inference cost entirely. The Gym-Anything paper tested this argument directly, and the results warrant careful reading.

The paper evaluated two open-weight models from the Qwen3-VL family: a 2B parameter base model and a 4B parameter base model. On CUA-World-Test under standard conditions, Qwen3-VL-2B achieves a 1.6% pass rate. Qwen3-VL-4B achieves 3.9%. The paper then distilled successful trajectories from Kimi-K 2.5 into the 2B model using CUA-World-Train data, producing a model that reaches 4.4% pass rate, outperforming the base 4B model. The authors present this as a meaningful result: a 2B distilled model beating a model twice its size. That framing is accurate on its own terms.

It does not address the more important comparison: 4.4% versus 12.8% for Kimi-K 2.5, and 22.6% for Gemini 3 Flash on the same benchmark.

The gap widens under conditions that matter most for enterprise deployment. The paper stratified results by visual complexity and domain knowledge. For small open-weight models, the findings are specific: Qwen3-VL-2B achieves a 3.2% pass rate on low-complexity software and 0.0% on high-complexity software. Distillation improves both numbers but does not change the shape of the curve. The paper states directly that visual complexity creates a disparity for small models that distillation alone does not resolve. For domain knowledge, small models show a decline approximately three times steeper than large models when moving from general to specialized software.

This matters for the white paper argument because enterprise software is disproportionately high-complexity and domain-specialized. The benchmark tasks that most resemble real professional work, radiology imaging analysis, ERP reconciliation, legal document management, are precisely the tasks where small open-weight models show the steepest performance decline. The cost argument in favor of open-weight models assumes that performance loss is acceptable or recoverable through fine-tuning. The paper's data suggests the loss is not recoverable at small model sizes, at least not with current distillation methods on current architectures.

There is a second finding that the open-source cost argument tends to overlook: the distillation result itself reveals an architecture dependency. The paper found that Kimi-K 2.5, as an open-source model providing full reasoning chains, produced the best student models despite being a weaker performer than Claude Opus 4.5 as a teacher. The paper's hypothesis is that transparent reasoning chains provide richer training signal than high-quality but opaque trajectories. If correct, this means the quality of open-weight distillation is currently bounded by the availability of open reasoning chains from teacher models, not just by the student model's parameter count.

That bound has nothing to do with cost. It is an architectural constraint on what the distillation pipeline can learn.

The open-source cost argument assumes that performance loss from smaller models is recoverable through fine-tuning. The CUA-World data shows the loss is steepest on the tasks that matter most for enterprise deployment: high visual complexity, specialized domain knowledge, long-horizon workflows. Distillation helps. It does not close the gap. And neither model class, proprietary or open-weight, addresses the underlying architecture problem. Both hold intent in the context window. Both are subject to context fatigue. The cost of the model does not change the cost of the architecture.
Cross-reference: Paper 12: "The Intent Economy" | Paper 14: "The Necessary Sequence"

Section 06Governed Operations Platforms and the Provenance Trap

A third category of proposed solution deserves separate treatment, because it uses the vocabulary of governance while preserving the architecture of detection. A growing class of agentic operations platforms claims to solve the long-horizon agent problem through what they variously describe as bounded autonomy, decision provenance, and composite AI governance. The claims are specific enough to evaluate, and the evaluation matters for understanding what a substrate solution actually requires.

The pattern is consistent across this category. An orchestration layer coordinates multiple specialized agents. Each agent's reasoning steps, tool calls, and actions are logged in a provenance layer that records input data, policy checks, and outcomes. When an agent's confidence falls below a threshold, it escalates to a human operator. The platform presents this combination as governed agentic execution.

Three architectural facts distinguish this pattern from a substrate solution, and each maps directly to a failure mode the Gym-Anything benchmark exposed.

First

The provenance layer logs what happened. It does not determine what is permitted to happen. The distinction is the one Paper 21 drew between attestation and governance: a record of behavior produced after the fact is not a governing constraint applied before execution. An agent in these platforms acts and then is logged. A substrate solution evaluates the action before it executes and records the determination, not the behavior.

Context fatigue produces premature task completion, placeholder substitution, and mid-trajectory abandonment precisely because nothing evaluated the proposed action against the original intent before it executed. Logging the action afterward confirms the failure. It does not prevent it.

Second

Confidence-threshold escalation is a detection mechanism, not a governance mechanism. The agent operates until its own internal confidence signal triggers a handoff. Two problems follow immediately. The agent's confidence signal is unreliable under context fatigue: the Gym-Anything paper shows that agents declare tasks complete prematurely, which is a high-confidence error, not a low-confidence one. An agent that does not know it has drifted from its original intent will not escalate, because it does not know it has a problem. The escalation architecture assumes the agent can accurately assess its own deviation. The benchmark evidence is that it cannot.

Third

Composite AI grounding, combining generative reasoning with symbolic AI, causal models, or first-principles physics, does not address the intent persistence problem. A better-grounded agent still holds the task specification inside its context window. Grounding improves the quality of individual reasoning steps. It does not protect the governing intent of the overall trajectory from the retrieval dynamics that produce context fatigue. The two problems are orthogonal. Grounding is a reasoning improvement. Substrate is an architectural position.

Logging every action is not governing every action. Escalating when confidence is low does not catch high-confidence drift. Grounding reasoning in physics models does not hold intent outside the context window. Each of these is a real improvement over a bare language model. None of them is a substrate solution. The difference is not a matter of degree. It is a matter of where in the architecture governance is applied: before execution, or after it.
Cross-reference: Paper 21: "The Missing Substrate" | Paper 3: "The Ornithopter Mistake"

Section 07The Benchmark Gap and the Governance Gap Are the Same Gap

The Gym-Anything paper was written as a contribution to agent benchmarking research. Its authors were not trying to describe an AI governance problem. But the gap it measures is the same gap this series has been describing from a different direction, and the convergence is worth stating explicitly.

Paper 213 argued that the emerging consensus on AI governance presupposes a technical substrate it never names. The consensus wants continuous, machine-verifiable audits of AI system behavior. It cannot produce them from a detection-based stack, because a detection-based stack produces logs, not governed records. The missing layer is the one that evaluates intent before execution and records each determination as it happens.

The Gym-Anything paper is measuring the same absence from the agent performance side. An agent that loses track of its intent mid-trajectory is not being governed. It is being detected, after the fact, by a pass/fail evaluator that checks whether the final state matches the task specification. The evaluator does not prevent drift. It measures drift that has already occurred. The 6–8% pass rates on long-horizon tasks are not a measurement of what agents can do. They are a measurement of how severely ungovemed execution degrades over long trajectories.

This is also the answer to a question the benchmark raises but does not answer: why does adding an audit agent at the end of the trajectory help? Because the audit agent is, for one additional invocation, doing what a governing substrate does continuously: comparing the current state of the work against the original intent and identifying the gap. The improvement from 11.5% to 14.0% is what one retrospective comparison is worth. A governing substrate that makes that comparison at every step, before each action, is worth the entire 88–94% that currently fails.

Paper 21 noted the live Synergy governance event on June 4, 2026 (provenance anchor ens:WIN7N340),4 in which a generative model proposed an intent-conflicting action and Synergy rejected it before execution. That event was a single-step demonstration of the architecture this paper is describing at the trajectory level. Governed determination at every step is not a proposal. It is an operational capability. The Gym-Anything benchmark provides, at scale, the empirical frame for why it is necessary.

The 6–8% pass rate on realistic long-horizon tasks is not an agent capability problem. It is a measurement of how far ungoverned execution drifts from its stated intent over a long trajectory. Governing the trajectory from the substrate changes that arithmetic. It does not improve retrieval of a buried specification. It removes the specification from the window where it can be buried.
Cross-reference: Paper 3: "The Ornithopter Mistake" | Paper 21: "The Missing Substrate" | Synergy governance event ens:WIN7N340, June 4, 2026

Section 08Conclusion: The Ceiling Is the Architecture

The Gym-Anything paper is a careful and significant piece of work. Its methodological contribution, using agents to build agent environments at scale, is genuinely novel. Its empirical finding, that realistic enterprise software tasks defeat every current frontier model on long horizons, is the most rigorous confirmation of the capability gap yet produced. Its diagnosis of context fatigue as the primary failure mode gives the problem a specific and testable name.

Where this paper parts from the Gym-Anything analysis is on the prescribed direction. Longer context windows, better training on long-horizon tasks, and test-time auditing are all real improvements. They are also improvements that work within the architecture that produces the failure. An agent trained to complete tasks inside a context window will always be subject to the retrieval dynamics that degrade long-horizon performance. The ceiling is not a training or capacity parameter. It is the architecture.

Intent-native computing does not improve retrieval of a buried task specification. It removes the task specification from the position where it can be buried. The governing intent lives in the substrate, evaluated before each action, recorded at each determination, and persistent across the entire trajectory regardless of length. Context fatigue cannot reach it because context has no access to it. The trajectory can extend for ten thousand steps or a hundred thousand. The governing representation of what the work is for does not degrade with the trajectory, because it was never inside the trajectory.

The CMU researchers found the ceiling. This series has been specifying what is on the other side of it.

Essence® Doctrine
The agent proposes.
The context window forgets.

Detection is not Determination.

Intent held in the substrate
cannot be lost in the middle of it.
Footnotes
1
Pranjal Aggarwal, Graham Neubig, Sean Welleck. "Gym-Anything: Turn any Software into an Agent Environment." arXiv:2604.06126 [cs.LG], April 7, 2026. arxiv.org/abs/2604.06126
2
Nelson F. Liu, Kevin Lin, John Hewitt, et al. "Lost in the Middle: How Language Models Use Long Contexts." Transactions of the Association for Computational Linguistics, 2024. Cited by the Gym-Anything paper as the basis for the context fatigue diagnosis. arxiv.org/abs/2307.03172. You may want to verify this URL against the primary source.
3
MindAptiv White Paper 21: "The Missing Substrate: Why the Emerging Consensus on AI Governance Presupposes an Architecture It Has Not Named." Ken Granville, MindAptiv, July 2026. mindaptiv.com/missing-substrate
4
Synergy governance event, provenance anchor ens:WIN7N340, June 4, 2026. Live documentation of a Synergy® rejection of a generative model's proposed action prior to execution. Internal MindAptiv record; no public URL.
5
CUA-World project page, CMU L3 Lab. Interactive paper and dataset access. cmu-l3.github.io/gym-anything
Sources & References
All figures cited from the Gym-Anything paper should be verified directly against the primary source. Model names are as they appear in the paper.
01
Pranjal Aggarwal, Graham Neubig, Sean Welleck. "Gym-Anything: Turn any Software into an Agent Environment." arXiv:2604.06126 [cs.LG], April 7, 2026. Primary source for all benchmark figures cited in this paper.
arxiv.org/abs/2604.06126
02
Nelson F. Liu, Kevin Lin, John Hewitt, et al. "Lost in the Middle: How Language Models Use Long Contexts." Transactions of the Association for Computational Linguistics, 2024. The "Lost in the Middle" retrieval degradation research cited by the Gym-Anything paper as the basis for the context fatigue diagnosis.
arxiv.org/abs/2307.03172
03
MindAptiv White Paper 21: "The Missing Substrate: Why the Emerging Consensus on AI Governance Presupposes an Architecture It Has Not Named." Ken Granville, MindAptiv, July 2026.
mindaptiv.com/missing-substrate
04
Synergy governance event, provenance anchor ens:WIN7N340, June 4, 2026. Live documentation of a Synergy® rejection of a generative model's proposed action prior to execution. Referenced in Paper 21 and cited here as operational evidence of pre-execution governed determination.
05
CUA-World project page, CMU L3 Lab. Interactive paper and dataset access.
cmu-l3.github.io/gym-anything
White Paper Series · The Governed Machine
1The Civilizational Fault Line 2We Are Building the Wrong Machine 3The Ornithopter Mistake 4The Convergence 5The Four Horsemen of the Knowledge Apocalypse 6What the Insiders Confirmed 7The Metaphor Trap 8The Recall Standard 9The $1 Trillion Governance Gap 10The Litigation Layer 11The Scale of Intent 12The Intent Economy 13The Session Illusion 14The Necessary Sequence 15The Wrong Race 16The Ledger That Is Intent-Driven 17The Agency Illusion 18The Substrate 19The End of the Mean 20Era 3: The Architecture of the Next Civilization 21The Missing Substrate 22The Context Fatigue Ceiling ← this paper 23The Iceberg Stays Frozen 24The Dependency Tax 25The Record That Was Never Kept 26Composable by Default 27Do No Harm 28The Stack Replacement Thesis 29The Moat Is the Code 30The Last Platform War 31Beyond the Agent: Intent-Native Execution 32The Hardware Imagination 33The Architecture Tax 34The Tokenization Ceiling 35The Payment Moment 36The Oracle Problem 37The Reviewer Problem 38The Provenance Fallacy 39Role Without Determination 40Known and Funded Anyway 41The Style Confusion Proof 42The Verification Tax 43The Pause Reflex 44The Human Margin 45The Balance of Power Fallacy 46The Liability Backstop 47One Substrate, Every Signal 48The Attribution Problem 49The Consciousness Ceiling 50The Detection Patch 51The Consumptive Machine 52The Agent That Isn't 53The Legibility Gap 54The Semiotic Machine 55The Transpilation Ceiling 56The Provisioning Ceiling 57The Reservation Ceiling 58The Circularity Ceiling 59The Coexistence Ceiling 60The Conformance Ceiling 61The Preservation Ceiling 62The Parity Clause 63The Governed Boundary 64The Transcript Problem 65The Unpaired System 66The Memory Ceiling 67The Admission Gap 68The Wrong Ask 69The Best Case 70The Last Chokepoint 71The Fourth Step 72The Adoption Standard 73The Same Weekend 74Sixty to One 75Coordinates, Not Correlations 76The Governability Axis 77Era 3, Confirmed 78The Eleventh Rule 79The Seventh Admission 80The Authorization Gap 81The Authorship Fallacy 82The Camera and the Vault 83Cleared to Proceed 84A Class, Not a Product 85The Inherited Playbook