1 2 3 4 5 6 7 8 9 10 11 12 13

The Instruction Set Boundary

How Deep Generation Goes, and Who Finishes the Job

A kernel compiled for one vendor's accelerator instruction set is not slower on another vendor's hardware. It is worth nothing. This is the most literal form the fixed commitment takes anywhere in this series, and it sits on the components that decide how fast the expensive work runs. The industry's answer is to stop short of the instruction set and let the vendor's driver finish, which is good engineering and puts a compiler nobody controls in the middle of every execution.

Ken Granville CEO & Co-Founder, MindAptiv Essence Paper 6 The Common Substrate August 2026
Substrate
GPUs and accelerators
Generating today
SPIR-V
Reconstituting
PTX, GCN
Status
SPIR-V shipping · rest roadmap
Abstract

Papers IV and V both ended by pointing at the same place. The managed runtime generates excellent instructions for a program written not to know what machine it is on, and the conformance floor guarantees behavior while saying nothing about the accelerator. The properties neither of them can describe are properties of components, and the accelerator is the component where the difference is largest and the money is. It is also where the fixed commitment is at its most literal: a kernel compiled for one vendor's instruction set does not degrade gracefully on another's, it does not run at all.

The industry's response was to stop short. Rather than shipping machine instructions, accelerator toolchains ship an intermediate representation and let the vendor's driver perform the final compilation on the target machine. That is deferred compilation, it is the correct instinct, and the accelerator world arrived at it well before the general software industry did. It also has a consequence that is rarely stated as a cost: the driver's compiler is now an unavoidable, opaque, independently versioned participant in every execution, expensive enough that its output has to be cached, and invalidated whenever the driver changes. This paper argues that the boundary, how far down your own generation reaches before somebody else's compiler takes over, is the thing worth reasoning about, that the industry's real lock-in lives at that boundary rather than in the silicon beneath it, and states plainly where this architecture's boundary currently sits: SPIR-V is generating today, with PTX and GCN reconstitution ahead rather than delivered.

Section 01Where the Commitment Is Most Literal

Throughout this series the fixed commitment has been costly but survivable. A binary compiled for the wrong Linux distribution usually fails at load with a message. A program written to a conformance floor runs everywhere and leaves capability unused. The costs are real and they are all recoverable.

The accelerator is different in kind. Instruction sets from different vendors, and often across generations from one vendor, are not dialects of a shared language. Compiled output for one is not a degraded input for another; it is not an input at all. There is no fallback path, no compatibility mode, no slower-but-working outcome. The artifact is either for this hardware or it is inert.

That would be an academic point if accelerators were peripheral, and they are the opposite. This is the hardware that determines whether the expensive workloads in the world finish in an hour or a week, and it is the component the industry is spending the most money on. The place where the commitment is most absolute is also the place where the stakes are highest, which is why it has attracted the most engineering and the most durable market power.

Section 02Everyone Stops at an Intermediate Representation

Faced with an artifact that is worthless on the wrong hardware, the accelerator world did something this series should acknowledge as correct: it stopped shipping the artifact.

What ships instead is an intermediate representation. Work is compiled part of the way (far enough to be checked, optimized and distributed, not so far as to be bound to a specific processor) and the final step to actual instructions happens on the target machine, performed by the vendor's driver, which is the only party that knows exactly what is installed. This is deferred compilation, arrived at for exactly the reason Paper 1 argues everyone should have arrived at it, and the accelerator world got there before the general software industry did. Paper 4 made the same observation about a managed runtime on the phone. Both are instances of the same good instinct: do not decide what you cannot yet know.

The result is that almost nobody in this part of the industry ships machine code, and almost everybody ships something one step above it. Which raises a question the industry rarely asks out loud, because the answer has been the same for so long that it stopped looking like a choice: how far down should your own generation reach before you hand off?

Section 03What Leaving the Vendor's Compiler in the Path Costs

Handing off at the intermediate representation puts a compiler you did not write, cannot inspect and do not ship into the execution path of everything you build. Four consequences follow, and none of them are hypothetical to anyone who has shipped accelerated software.

It is opaque

The transformation from your representation to the instructions that actually execute happens inside a component you cannot read. When output is slower than expected, the reason is on the other side of a boundary you cannot cross, and the available response is to perturb the input and observe. That is not engineering, it is divination with a profiler.

It is versioned by somebody else

The compiler changes when the driver changes, on a schedule set by the vendor rather than by you, and the same input can produce materially different output before and after. Software that was fast last month can be slower this month with no change on your side, and the change was not announced because from the vendor's perspective nothing user-visible happened.

It is expensive enough to be visible

Final compilation is costly enough that it cannot be done casually at the moment work is submitted, which is why the industry caches results aggressively and why users experience the failure of that caching directly, as the pause the first time something new appears on screen. An entire body of engineering practice exists to pre-warm those caches, and it exists because the compile is in the wrong place.

The cache is invalidated by the thing you do not control

A driver update discards the accumulated results, and the cost is paid again. Paper 9 described the update as the most dangerous routine operation at the edge; here it is merely the most expensive, and it is triggered by a party with no visibility into what it costs you.

The Structural Point
Deferring the final compile is right. Deferring it into somebody else's opaque, independently versioned compiler is a different decision that arrived attached to the first one, and the industry has largely stopped noticing that they are separable.

Section 04The Lock-In Is the Toolchain, Not the Silicon

This section is the one an investor should read, because it explains a market structure that is usually attributed to the wrong cause.

The dominant position in accelerated computing is routinely described as a hardware advantage. Hardware advantages are real and they are also transient: competitors ship parts with comparable characteristics, and buyers who care about cost notice. The reason the position persists through those cycles is that the switching cost is not in the hardware. It is in everything written against one vendor's toolchain (the source, the libraries, the kernels, the accumulated tuning, the people who know it), none of which transfers, because it was written in a vendor-specific dialect and compiled through a vendor-specific path.

Put precisely: the moat is the boundary described in Section 02. Because everyone stops at the vendor's representation, everyone's work is expressed in terms of one vendor's toolchain, and the accumulated body of that work is what makes leaving expensive. The silicon is substitutable and the corpus is not.

The structural consequence for an architecture that generates instructions from declared intent is worth stating carefully, because it is easy to overstate and this paper is not going to. Work expressed as intent is not written in any vendor's dialect, so there is no corpus to port; targeting different hardware is a backend rather than a rewrite. That does not mean the hardware becomes trivially reachable. Generating for an instruction set requires knowing that instruction set, vendors differ enormously in how much they document, and a backend is real engineering. It means the cost of supporting another vendor is borne once, by the platform, instead of being borne repeatedly by every organization that wrote code in the first vendor's terms.

Section 05Where This Architecture Sits Today

The honest statement of the current position is short, and the substrate strip at the top of this paper carries it. SPIR-V is generating today. PTX and GCN were generated previously and are being reconstituted, which places them ahead rather than delivered, and no reader should take anything in this paper as a claim that they are available now.

What that supports and what it does not are both worth being clear about. SPIR-V is a broad target and reaching it means work declared as intent becomes accelerator instructions through this architecture's own generation rather than through source written in a vendor's language, which is the point of Section 04: there is no vendor-dialect corpus accumulating. It also means that, at this target, the final step into the specific processor's instructions is still completed on the machine by the vendor's component, with the properties Section 03 describes. The boundary is where it is, and moving it further down is what the reconstitution work is for.

This is also the layer Papers IV and V were pointing at when each of them stopped. The managed runtime cannot generate for an accelerator the program had no vocabulary to describe, and the conformance floor does not tell an application which graphics architecture it received. Both gaps are addressed at the same place, which is generation against the component rather than against an abstraction of it.

The emission depth ladder A chain from source through intermediate representation, then the vendor driver compiler, then native GPU instructions. Conventional toolchains stop at the intermediate representation and the vendor driver completes the compilation. This architecture generates SPIR-V today, with vendor-native targets reconstituting. THE EMISSION DEPTH LADDER SOURCEshader or kernel IRSPIR-V, PTX DRIVER COMPILERvendor's, opaque NATIVE GPU ISAwhat runs CONVENTIONAL TOOLCHAINS STOP HERE WHERE THE INSTRUCTIONS ACTUALLY ARE GENERATING TODAY: SPIR-V RECONSTITUTING: PTX AND GCN · NOT AVAILABLE NOW
Figure 1. The ladder is the paper. Every toolchain picks a rung to stop on, and everything below that rung is performed by whoever owns the next component. The question is never whether to defer the final compile, which is correct, but who is holding the compiler when it happens.

Section 06What This Paper Does Not Claim

PTX and GCN are not available. They were generated previously and are being reconstituted. Only SPIR-V is generating today, and any reading of this paper that treats the vendor-native targets as current is a misreading the substrate strip and Section 05 both try to prevent.

No comparison against vendor toolchains has been run. No benchmark is reported here measuring generated output against what a vendor's own compiler produces from equivalent work, on any hardware. Section 03 describes structural properties of leaving a compiler in the path, not a measured performance deficit.

Vendor toolchains are not badly engineered. They are, in general, extremely good, written by people with information nobody outside the company has. Section 03's argument is about where the compiler sits and who controls it, not about its quality.

Not all hardware is equally reachable. Generating for an instruction set requires documentation of that instruction set, and vendors differ substantially in what they publish. A backend is real engineering with a real cost, and the claim in Section 04 is that the cost is borne once by the platform rather than repeatedly by every customer, not that it is small.

The series' performance figures are not GPU-backend figures. The acceleration and energy results cited elsewhere were measured on specific workloads under stated conditions with AWS and Rowan University's Digital Engineering Hub. They are not attributable to accelerator instruction generation in general.

Series context · This paper reports no engagement, evaluation or partnership with any accelerator vendor, and names instruction set architectures only as technical targets.

Section 07Where the Series Goes Next

The second movement continues one layer down. Paper 7 takes the processor architectures themselves, where the same commitment produces a different and more visible failure: an architecture transition, in which an entire software ecosystem discovers simultaneously that everything it shipped was compiled for a processor family it is now leaving. The industry has just finished doing this once and is being asked to do it again.

The position there is the same shape as the one stated here, and just as bounded: x86-64 is what the shipping binary targets today, and the other architectures are roadmap rather than delivered.

What This Paper Adds to the Series
Every earlier paper asked when the instruction decision gets made. This one asks how far it goes and who owns the last step, and finds that the answer explains a market structure usually credited to the silicon. The corpus written in a vendor's dialect is the moat. Work expressed as intent does not accumulate one.
The Common Substrate: Essence Paper 6

Deferring the final compile is correct.
Deferring it into a compiler you do not control is a separate decision.

The accelerator is where a compiled artifact is not merely wrong for other hardware but worthless on it, and the industry's answer (ship an intermediate representation, let the driver finish) is the right instinct arrived at early. It also placed an opaque, independently versioned compiler in the path of every execution, expensive enough to require caching and invalidated by updates nobody consults you about. And because everyone stops at the vendor's representation, everyone's accumulated work is written in one vendor's terms, which is the actual switching cost that the hardware gets credited for. SPIR-V is generating today. PTX and GCN are reconstituting. Both of those are stated as what they are.

The Common Substrate → Request Platform Access

White Paper Series · The Common Substrate