Why Frontier Agents Fail on Long-Horizon Work, and Why Scaling Cannot Fix It
Carnegie Mellon University's Gym-Anything paper, published April 2026,1 is the most rigorous large-scale measurement of computer-use agents on realistic enterprise software to date. Its CUA-World benchmark spans 10,215 tasks across 200 applications covering all 22 major U.S. occupation groups.
The headline result is unambiguous: on long-horizon tasks, every frontier model tested achieved a pass rate below 8% under standard cost constraints. The paper identifies the primary failure mode and gives it a name: context fatigue. After processing hundreds of thousands of tokens, agents lose track of what they were supposed to accomplish. They declare tasks complete prematurely, substitute placeholder data for real outputs, and abandon their original intent mid-trajectory.
The authors treat this as a technical problem to be solved with better training, longer context windows, and test-time auditing. This paper argues the diagnosis is correct and the prescription is incomplete.
Context fatigue is not a memory problem. It is an architecture problem. An agent that holds intent inside its context window will always be vulnerable to losing it. A platform that holds intent outside the model, as a governed, persistent representation evaluated before execution, cannot forget what the work is for. The fix is not a longer window. It is a different substrate.
The hard problem in AI agent research has always been measurement. Existing benchmarks test narrow, short tasks: click a button, fill a form, navigate a website, configure a system setting. These are useful calibration tools for individual capabilities. They are a poor proxy for whether an agent can do real work.
The CMU team's response to this problem is methodologically significant.5 Rather than handcrafting a fixed set of tasks, they used agents to build the benchmark itself. A coding agent writes setup scripts, installs software, loads real-world data, and configures each application to a realistic starting state. An independent audit agent then verifies the setup against a quality checklist, issuing corrections before any task is attempted. The result is CUA-World: 10,215 tasks across 200 applications, grounded in a taxonomy derived from U.S. GDP data by occupation group, spanning domains from medical imaging analysis to enterprise resource planning to legal document management.1
This is the first benchmark that approximates the shape of actual professional work. Not a simplified proxy. Not a curated toy environment. Software that real workers use, populated with realistic data, requiring multi-step workflows that resemble what the work actually demands.
The results from that benchmark are reproduced in the table below, drawn directly from the paper. All figures cited here should be verified against the primary source at arxiv.org/abs/2604.06126.1
| Model | Type | Benchmark | Steps / Cap | Pass Rate |
|---|---|---|---|---|
| CUA-WORLD-TEST (Standard Benchmark) | ||||
| Gemini 3 Flash | Proprietary | CUA-World-Test | 200 steps | 22.6% |
| Kimi-K 2.5 | Open-source | CUA-World-Test | 200 steps | 12.8% |
| Qwen3-VL-4B (base) | Open-weight | CUA-World-Test | 200 steps | 3.9% |
| Qwen3-VL-2B (distilled) | Open-weight | CUA-World-Test | 200 steps | 4.4% |
| Qwen3-VL-2B (base) | Open-weight | CUA-World-Test | 200 steps | 1.6% |
| CUA-WORLD-LONG (Long-Horizon Benchmark) | ||||
| Gemini 3 Flash | Proprietary | CUA-World-Long | 500 steps / $5 | 7.5% |
| Kimi-K 2.5 | Open-source | CUA-World-Long | 500 steps / $5 | 5.5% |
| Claude Sonnet 4.6 | Proprietary | CUA-World-Long | 500 steps / $5 | 6.0% |
| GPT-5.4 | Proprietary | CUA-World-Long | 500 steps / $5 | 3.0% |
| GPT-5.4 | Proprietary | CUA-World-Long | 2,000 steps / uncapped | 27.5% |
| Gemini 3 Flash | Proprietary | CUA-World-Long | 2,000 steps / uncapped | 11.5% |
| Gemini 3 Flash + TTA | Proprietary | CUA-World-Long | 2,000 steps / uncapped | 14.0% |
Two numbers in this table require additional attention because they reframe the others. GPT-5.4's capped pass rate of 3.0% is not primarily a capability failure. The paper notes the model exhausts its $5 cost budget in roughly 150 steps, well below the 500-step ceiling. Its uncapped result of 27.5% at 2,000 steps is the strongest in the study. The cost cap is a real-world constraint, not an artifact of the benchmark, and it reveals something important: the most capable model is also the most expensive to run on long tasks. Neither raw capability nor cost efficiency is sufficient.
The second number that reframes the table is the Test-Time Auditing result. Adding an external audit agent to Gemini 3 Flash improves its uncapped long-horizon pass rate from 11.5% to 14.0%. The paper presents this as a practical improvement. It is also an inadvertent structural finding, and Section 03 takes it up directly.
The Gym-Anything paper does not simply report the pass rates. It diagnoses the failure. The authors examined agent trajectories on failed long-horizon tasks and identified a consistent pattern they call context fatigue, citing earlier research on the phenomenon under the label "Lost in the Middle."2
The description is precise. After processing hundreds of thousands of tokens across a long trajectory, the agent loses track of what it was originally trying to accomplish. The degradation is not random. It takes characteristic forms.
The agent declares the task complete while the application is at the wrong screen. It substitutes placeholder or sample data for the real outputs the task required. It completes a subset of the required steps and stops, without recognizing that the remaining steps were part of the same task. In each case, the original intent is present somewhere in the context window. The agent simply can no longer act on it reliably.
The paper speculates on the mechanism, consistent with the "Lost in the Middle" literature: information in the middle of a very long context is retrieved less reliably than information near the beginning or end. As a trajectory grows, the original task specification migrates toward the middle, and the agent's ability to act on it degrades. The longer the task, the more severe the degradation.
This distinction matters because the natural response to context fatigue is to make the context window longer. If the agent loses intent because the original task specification slides toward the middle of a long window, extend the window so the specification stays closer to the boundary. This is the prescription the scaling paradigm reaches for instinctively. It is also, as Section 04 argues, the wrong prescription, because it treats a structural problem as a capacity problem.
The most consequential result in the Gym-Anything paper is not the pass rates. It is the Test-Time Auditing experiment, and what it reveals about the architecture the authors were not trying to describe.
The setup is straightforward. After a primary agent completes its trajectory, a second model reviews the trajectory against the original task specification, identifies what was missed or done incorrectly, and feeds that assessment back. Under the 2,000-step uncapped condition, Gemini 3 Flash improves from 11.5% to 14.0% pass rate with this addition. The authors present this as a practical technique for improving agent performance at test time.
But notice what the technique actually does. It externalizes the memory of the original intent. The primary agent, having processed a long trajectory, has lost reliable access to what the task required. The audit agent, reading the trajectory fresh against the original specification, re-introduces the intent from outside the primary agent's degraded context. The improvement in pass rate is a direct measurement of how much intent the primary agent had lost by the end of its trajectory.
This is the same structural move Paper 213 identified in the AI governance debate. Just as the entity-governance proposal reaches for continuous, AI-facilitated auditing because periodic human audits are insufficient, the Gym-Anything paper reaches for an external audit agent because the primary agent's internal state is insufficient. In both cases, the prescription is a governance layer added after the fact. In both cases, the underlying diagnosis points at the same missing architecture: a substrate that holds intent persistently, outside the executing model, from the beginning of the task rather than as a correction at the end.
The scaling response to context fatigue is not irrational. If the agent loses intent because the task specification is buried in a long context, a longer context window keeps the specification closer to retrievable positions. A better-trained model retrieves middle-context information more reliably. Both improvements are real and measurable. The uncapped GPT-5.4 result of 27.5% is evidence that removing cost and step constraints substantially improves performance. Progress on context length and retrieval fidelity will continue.
But the progress has a ceiling, and the ceiling is structural. Here is why.
A model holding intent inside its context window must carry that intent through every step of the trajectory. Every tool call, every screenshot observation, every intermediate result adds tokens to the window. The task specification, fixed at the beginning, does not grow; everything else does. The ratio of task-specification tokens to total context tokens decreases monotonically as the trajectory extends. No context length eliminates this dynamic; it only shifts the point at which degradation becomes severe. The "Lost in the Middle" problem2 is not a bug to be patched. It is a property of the architecture.
There is a second, independent ceiling: cost. The uncapped results in the Gym-Anything paper are not deployment conditions. They are laboratory measurements. At 2,000 steps per task, the per-task inference cost for frontier models is not commercially viable for the class of work CUA-World tests. The gap between the capped pass rates and the uncapped pass rates is not a benchmark artifact. It is a preview of the real constraint that enterprise deployment will impose. More capable models are more expensive per step. Long-horizon tasks require more steps. Both trends move in the direction of higher cost, not lower.
Intent-native architecture addresses both ceilings simultaneously, because it moves intent out of the context window entirely. In the Essence® platform, the wantverse holds what the user or system wants to accomplish as a governed, persistent representation that exists independently of any individual model invocation. Synergy® evaluates each proposed action against that representation before execution. The model does not need to remember the original task specification in its context window, because the task specification is not in the context window. It is in the substrate.
A 2,000-step trajectory does not erode the governing intent, because the governing intent was never inside the trajectory to begin with.
The user or calling system expresses what is to be accomplished. This declaration is stored as a governed representation in the wantverse, not as tokens in a model's context window.
Before any step executes, Synergy® evaluates the proposed action against the persistent intent record. The model proposes. The substrate determines whether the proposal is consistent with what the task actually requires.
Context fatigue is a failure of retrieval: the agent cannot reliably access intent buried in a long window. Intent held outside the window cannot be buried in it. The governing representation is not subject to the "Lost in the Middle" dynamic, because it is not in the middle of anything.
Each evaluated action produces an AptivRecord: what was intended, what was permitted, what executed, and under what constraint. The trajectory produces a governed ledger, not a log. Audit, if required, verifies the ledger rather than reconstructing intent from a long context after the fact.
The Gym-Anything paper shows that Test-Time Auditing recovers some of the intent the primary agent lost. That recovery comes at the cost of an additional model invocation at the end of a trajectory that has already consumed thousands of steps. The intent-native alternative is not a correction applied at the end. It is a governing constraint applied at every step. The difference is not marginal. It is the difference between fixing drift after it has accumulated and making drift architecturally impossible.
A predictable response to the cost ceiling documented above is to substitute smaller, open-weight models for frontier proprietary ones. If GPT-5.4 costs approximately $18 per long-horizon trajectory uncapped, and Gemini 3 Flash costs approximately $16, the argument runs: use a fine-tuned open-weight model instead, run it locally, and eliminate the per-step inference cost entirely. The Gym-Anything paper tested this argument directly, and the results warrant careful reading.
The paper evaluated two open-weight models from the Qwen3-VL family: a 2B parameter base model and a 4B parameter base model. On CUA-World-Test under standard conditions, Qwen3-VL-2B achieves a 1.6% pass rate. Qwen3-VL-4B achieves 3.9%. The paper then distilled successful trajectories from Kimi-K 2.5 into the 2B model using CUA-World-Train data, producing a model that reaches 4.4% pass rate, outperforming the base 4B model. The authors present this as a meaningful result: a 2B distilled model beating a model twice its size. That framing is accurate on its own terms.
It does not address the more important comparison: 4.4% versus 12.8% for Kimi-K 2.5, and 22.6% for Gemini 3 Flash on the same benchmark.
The gap widens under conditions that matter most for enterprise deployment. The paper stratified results by visual complexity and domain knowledge. For small open-weight models, the findings are specific: Qwen3-VL-2B achieves a 3.2% pass rate on low-complexity software and 0.0% on high-complexity software. Distillation improves both numbers but does not change the shape of the curve. The paper states directly that visual complexity creates a disparity for small models that distillation alone does not resolve. For domain knowledge, small models show a decline approximately three times steeper than large models when moving from general to specialized software.
This matters for the white paper argument because enterprise software is disproportionately high-complexity and domain-specialized. The benchmark tasks that most resemble real professional work, radiology imaging analysis, ERP reconciliation, legal document management, are precisely the tasks where small open-weight models show the steepest performance decline. The cost argument in favor of open-weight models assumes that performance loss is acceptable or recoverable through fine-tuning. The paper's data suggests the loss is not recoverable at small model sizes, at least not with current distillation methods on current architectures.
There is a second finding that the open-source cost argument tends to overlook: the distillation result itself reveals an architecture dependency. The paper found that Kimi-K 2.5, as an open-source model providing full reasoning chains, produced the best student models despite being a weaker performer than Claude Opus 4.5 as a teacher. The paper's hypothesis is that transparent reasoning chains provide richer training signal than high-quality but opaque trajectories. If correct, this means the quality of open-weight distillation is currently bounded by the availability of open reasoning chains from teacher models, not just by the student model's parameter count.
That bound has nothing to do with cost. It is an architectural constraint on what the distillation pipeline can learn.
A third category of proposed solution deserves separate treatment, because it uses the vocabulary of governance while preserving the architecture of detection. A growing class of agentic operations platforms claims to solve the long-horizon agent problem through what they variously describe as bounded autonomy, decision provenance, and composite AI governance. The claims are specific enough to evaluate, and the evaluation matters for understanding what a substrate solution actually requires.
The pattern is consistent across this category. An orchestration layer coordinates multiple specialized agents. Each agent's reasoning steps, tool calls, and actions are logged in a provenance layer that records input data, policy checks, and outcomes. When an agent's confidence falls below a threshold, it escalates to a human operator. The platform presents this combination as governed agentic execution.
Three architectural facts distinguish this pattern from a substrate solution, and each maps directly to a failure mode the Gym-Anything benchmark exposed.
The provenance layer logs what happened. It does not determine what is permitted to happen. The distinction is the one Paper 21 drew between attestation and governance: a record of behavior produced after the fact is not a governing constraint applied before execution. An agent in these platforms acts and then is logged. A substrate solution evaluates the action before it executes and records the determination, not the behavior.
Context fatigue produces premature task completion, placeholder substitution, and mid-trajectory abandonment precisely because nothing evaluated the proposed action against the original intent before it executed. Logging the action afterward confirms the failure. It does not prevent it.
Confidence-threshold escalation is a detection mechanism, not a governance mechanism. The agent operates until its own internal confidence signal triggers a handoff. Two problems follow immediately. The agent's confidence signal is unreliable under context fatigue: the Gym-Anything paper shows that agents declare tasks complete prematurely, which is a high-confidence error, not a low-confidence one. An agent that does not know it has drifted from its original intent will not escalate, because it does not know it has a problem. The escalation architecture assumes the agent can accurately assess its own deviation. The benchmark evidence is that it cannot.
Composite AI grounding, combining generative reasoning with symbolic AI, causal models, or first-principles physics, does not address the intent persistence problem. A better-grounded agent still holds the task specification inside its context window. Grounding improves the quality of individual reasoning steps. It does not protect the governing intent of the overall trajectory from the retrieval dynamics that produce context fatigue. The two problems are orthogonal. Grounding is a reasoning improvement. Substrate is an architectural position.
The Gym-Anything paper was written as a contribution to agent benchmarking research. Its authors were not trying to describe an AI governance problem. But the gap it measures is the same gap this series has been describing from a different direction, and the convergence is worth stating explicitly.
Paper 213 argued that the emerging consensus on AI governance presupposes a technical substrate it never names. The consensus wants continuous, machine-verifiable audits of AI system behavior. It cannot produce them from a detection-based stack, because a detection-based stack produces logs, not governed records. The missing layer is the one that evaluates intent before execution and records each determination as it happens.
The Gym-Anything paper is measuring the same absence from the agent performance side. An agent that loses track of its intent mid-trajectory is not being governed. It is being detected, after the fact, by a pass/fail evaluator that checks whether the final state matches the task specification. The evaluator does not prevent drift. It measures drift that has already occurred. The 6–8% pass rates on long-horizon tasks are not a measurement of what agents can do. They are a measurement of how severely ungovemed execution degrades over long trajectories.
This is also the answer to a question the benchmark raises but does not answer: why does adding an audit agent at the end of the trajectory help? Because the audit agent is, for one additional invocation, doing what a governing substrate does continuously: comparing the current state of the work against the original intent and identifying the gap. The improvement from 11.5% to 14.0% is what one retrospective comparison is worth. A governing substrate that makes that comparison at every step, before each action, is worth the entire 88–94% that currently fails.
Paper 21 noted the live Synergy governance event on June 4, 2026 (provenance anchor ens:WIN7N340),4 in which a generative model proposed an intent-conflicting action and Synergy rejected it before execution. That event was a single-step demonstration of the architecture this paper is describing at the trajectory level. Governed determination at every step is not a proposal. It is an operational capability. The Gym-Anything benchmark provides, at scale, the empirical frame for why it is necessary.
The Gym-Anything paper is a careful and significant piece of work. Its methodological contribution, using agents to build agent environments at scale, is genuinely novel. Its empirical finding, that realistic enterprise software tasks defeat every current frontier model on long horizons, is the most rigorous confirmation of the capability gap yet produced. Its diagnosis of context fatigue as the primary failure mode gives the problem a specific and testable name.
Where this paper parts from the Gym-Anything analysis is on the prescribed direction. Longer context windows, better training on long-horizon tasks, and test-time auditing are all real improvements. They are also improvements that work within the architecture that produces the failure. An agent trained to complete tasks inside a context window will always be subject to the retrieval dynamics that degrade long-horizon performance. The ceiling is not a training or capacity parameter. It is the architecture.
Intent-native computing does not improve retrieval of a buried task specification. It removes the task specification from the position where it can be buried. The governing intent lives in the substrate, evaluated before each action, recorded at each determination, and persistent across the entire trajectory regardless of length. Context fatigue cannot reach it because context has no access to it. The trajectory can extend for ten thousand steps or a hundred thousand. The governing representation of what the work is for does not degrade with the trajectory, because it was never inside the trajectory.
The CMU researchers found the ceiling. This series has been specifying what is on the other side of it.