Technical Architecture · Chameleon® / Wantware

Composite Job Designs
are not benchmarks.

Every computing benchmark assumes a fixed workload class. Composite Job Designs operate across the full range of what computing can do, and they adapt continuously as hardware, conditions, and objectives change.

Published June 2026 Platform Wantware / Chameleon® Audience Infrastructure · AI Platform · Engineering
Conventional benchmarks
Measure one workload type on fixed hardware at a point in time. Results describe past performance of a specific executable. They cannot tell you what the same hardware can do on a different problem, or what it could do if execution were restructured.
Composite Job Designs
Express what the system must accomplish, not how. Wantware synthesizes the instruction sequences at runtime, continuously, against observed hardware behavior. The CJD spans any workload class computing can execute.
01 · The Benchmark Problem

Benchmarks measure a fixed past.
CJDs govern a dynamic future.

A benchmark is a contract between a workload and a moment in time. SPEC CPU measures integer and floating-point performance on a defined suite. MLPerf measures inference throughput on a defined model. LINPACK measures dense linear algebra. Each one is useful, and each one is irrelevant the moment the workload changes, the hardware changes, or the objective changes.

This irrelevance is structural, not incidental. Benchmarks are designed to hold everything fixed so that hardware can be compared. That constraint is exactly what makes them poor descriptions of production computing, where workloads are heterogeneous, hardware is multi-vendor, and performance objectives shift continuously.

"A benchmark tells you how fast your car goes on one track. A Composite Job Design tells the car how to drive any road it encounters, and rewrites the transmission map in real time."

Composite Job Designs are the alternative design primitive. They do not measure; they govern. A CJD declares the classes of work the system performs: what operations are in scope, what hardware behavior to optimize against, what constraints define success. The execution itself (the machine instructions, the memory access patterns, the kernel boundaries and scheduling) is generated, not frozen.

Why this distinction matters for AI workloads
AI inference and training are not single workloads. A single model forward pass involves memory transfers, arithmetic on dense tensors, softmax normalization, attention scoring, and output projection, each with different memory-bandwidth characteristics, different arithmetic intensity, and different hardware bottlenecks. A benchmark that captures one of these well captures all of them poorly. A CJD covers all of them in one structure, and wantware optimizes each element against the hardware it is actually running on.

02 · Architecture

What a Composite Job Design
actually is.

A CJD is a runtime structure, not source code, not a kernel, not a configuration file. It describes the kinds of work the system must perform during execution and optimization. Wantware reads the CJD and generates the machine instructions that realize those kinds of work on the target hardware.

The separation is fundamental: the CJD is the control plane. Instruction synthesis is the output plane. Keeping them separate is what allows execution to adapt without recompilation.

Intent
Work class
Morpheus
Silicon
CJD execution model · intent to silicon

What a CJD contains

CJDs are composites: they describe multiple classes of work that may execute in parallel, in sequence, or conditionally based on observed hardware state. A single CJD for an AI inference workload might include:

Example CJD scope: AI inference pipeline
Memory transfer class
Governs how weights and activations move between host and device memory. Wantware restructures access patterns dynamically against observed latency and bandwidth.
Arithmetic execution class
Covers dense matrix operations, attention computation, normalization. Instruction sequences are synthesized to match the GPU's actual arithmetic throughput, not the compiler's assumptions.
Scheduling and synchronization class
Controls execution ordering, kernel launch parameters, and synchronization barriers. Eliminates idle gaps that conventional orchestration layers introduce.

What a CJD is not

A CJD is not a kernel. Kernels are fixed instruction sequences tuned for one workload at one point in time. A CJD is not a configuration file. It does not tell wantware how to execute; it tells wantware what to accomplish. And a CJD is not a benchmark: it does not measure; it governs.

This distinction matters because the value of a CJD is not in what it reports. It is in what it enables: execution that adapts continuously to the hardware it runs on, without recompilation, without framework dependencies, without manual retuning.

Does it cover what highly skilled engineers do by hand?

Yes, and it does it continuously, across the full chip ecosystem, without a human in the loop.

Expert performance engineers spend days or weeks on a single workload doing things like: avoiding cache misses through prefetch scheduling and access reordering, managing memory at the instruction level, reordering data to maximize locality and eliminate false sharing, orchestrating work across execution units to eliminate pipeline stalls and idle gaps, and structuring parallel execution to be lockless, removing mutex contention and semaphore deadlock conditions that are among the hardest failure modes in concurrent software.

Each of these techniques is hardware-specific. Results are valid only for that workload, on that silicon, at that moment. When the hardware changes, or when thermal throttling shifts the memory-to-compute ratio, or a co-located job competes for cache, the tuning is no longer optimal. The engineer starts over. Morpheus does not.

What Morpheus does in real time, continuously
Cache and memory management
Cache misses are avoided through prefetch scheduling and access-pattern restructuring. Memory is managed at the instruction level: allocation, locality, and movement adapted continuously to observed latency and bandwidth rather than compiler assumptions.
Data reordering
Data is reordered in flight to maximize memory locality, eliminate false sharing, and align access sequences to hardware cache-line geometry. This is not a one-time layout pass; it adapts as execution conditions change.
Work orchestration
Work is orchestrated across execution units: kernel boundaries, thread block sizing, scheduling order, and synchronization points are all restructured dynamically to eliminate idle gaps and pipeline stalls that fixed orchestration layers cannot see.
Lockless parallelism
Parallel execution is structured to eliminate mutex contention and semaphore deadlock conditions entirely. Morpheus synthesizes instruction sequences that coordinate concurrent work without locks, removing a class of correctness and performance failure that plagues hand-tuned parallel code.
Instruction-level restructuring
Loop unrolling, SIMD vectorization, instruction reordering, register allocation, pipeline stall removal, and branch prediction tuning, applied continuously against observed hardware behavior, not compiler heuristics frozen at build time.
Execution structure optimization
Kernel fusion, thread coarsening, warp divergence elimination, occupancy tuning, and latency hiding, without a recompile, without a framework change, and without a performance engineer in the loop.
Network protocol generation · Bus level management · Cumulative computing
Morpheus generates network protocol logic at runtime, not as static stack code but as synthesized instruction sequences matched to observed bus and interconnect behavior. Bus-level management means Morpheus governs how work and data move across PCIe, NVLink, memory buses, and interconnects directly, without an abstraction layer between intent and the physical bus. Cumulative computing means execution results compound: prior computation informs and shapes subsequent synthesis decisions, so the system grows more efficient over time rather than resetting to a fixed baseline on each invocation. These three capabilities are covered by MindAptiv's issued U.S. patents on the Morpheus method.

The difference is not just speed. A skilled engineer optimizes once and freezes the result. Morpheus optimizes continuously: every cycle, against the hardware state that actually exists at that moment. Lockless parallelism is not a tuning parameter; it is a structural property of how Morpheus synthesizes concurrent execution. Mutex contention and semaphore deadlock are not failure modes that get managed; they are eliminated at the instruction synthesis level.

Beyond the execution layer, Morpheus covers three capabilities that hand-tuning cannot reach at all: network protocol generation at runtime rather than at compile time; bus-level management that governs how work moves across PCIe, NVLink, and memory interconnects without an abstraction layer; and cumulative computing, where each invocation builds on prior results rather than starting from a fixed baseline. These capabilities, along with the core Morpheus synthesis method, are covered by MindAptiv's three issued U.S. patents, with no blocking prior art identified.

This is what the validated results reflect. The 20–114× acceleration range is not the product of a one-time tuning pass. It is the product of execution that never stops exploring the viable instruction space on the underlying hardware (across NVIDIA, AMD, Intel, Apple, Qualcomm, Broadcom, Marvell, and ARM silicon) within the constraints declared by the CJD.

The .wv format

CJD output is delivered in the .wv format: a purpose-built bitstream container designed around three properties that conventional formats cannot simultaneously satisfy.

Illustrative animation. Cipher names and counters shown are examples of the visual concept, not the production algorithm set or live measurements.

Illustrative animation. Cipher names and counters shown are examples of the visual concept, not the production algorithm set or live measurements.

Illustrative animation. Cipher names and counters shown are examples of the visual concept, not the production algorithm set or live measurements.

Illustrative animation. Cipher names and counters shown are examples of the visual concept, not the production algorithm set or live measurements.

.wv bitstream · illustrative
Bytes packed
0
Encryption rotations
0
Fragmentation
0%
Active cipher
n/a
.wv · Wantware bitstream format
Zero memory fragmentation
The bitstream is laid out to eliminate heap fragmentation entirely. Memory is contiguous by design: no allocator overhead, no gap accumulation, no compaction pauses at runtime.
Headerless compaction
There are no headers. The bitstream carries no framing overhead, no format negotiation, no version fields. The payload is the entirety of the stream, maximally compact, immediately decodable.
StreamWeave® encryption
Every .wv stream is encrypted in transit and at rest by StreamWeave®, MindAptiv's adaptive encryption layer that shifts algorithm combinations continuously to outpace threats that don't exist yet.

These three properties are not independent features; they are a single design decision. A headerless format has no anchor points for fragmentation to accumulate around. A contiguous bitstream has no inter-segment gaps for an adversary to exploit structural patterns. StreamWeave encryption applied to a compact, headerless stream yields a protected artifact with no format scaffolding exposed to analysis. The format is the security posture.

For infrastructure operators, this means CJD artifacts move between systems (across hyperscalers, edge nodes, and heterogeneous fleets) without format translation, without re-encryption handoffs, and without the memory overhead that conventional container formats impose at ingestion.


03 · Scope

Across the full range
of what computing does.

The phrase "the range of things computing can do" is not marketing language. It is a structural claim about CJD design. Conventional benchmarks are workload-specific by construction. CJDs are workload-class-agnostic by construction: the same CJD architecture that governs a GPU rendering pipeline can govern a signal-processing chain, a financial simulation, or a satellite telemetry stream.

This breadth is validated by Chameleon®'s measured results across workload types that have almost nothing in common at the application layer:

Light & energy propagation

Memory-traversal-dominant. SCE Lightwires: 125.5 ms → 1.1 ms (114× faster, 99.7% energy reduction). Stresses sustained-load memory access patterns.

Recursive fractal geometry

Arithmetic-dense. SCE Rust Mandala: 107.8 ms → 1.0 ms (108× faster). Demonstrates that compute-bound workloads benefit as much as memory-bound ones.

Signed distance field lighting

Same CJD, four hardware configurations. AMD Radeon (84×), Nvidia GPU on AWS (54×), Nvidia GPU on GCP (49×), Nvidia GPU on OCI (41×). No recompile between vendors.

Volumetric field simulation

Dense 2D data arrays. Red Nebula: 150.0 ms → 2.3 ms (65× faster). Representative of scientific and engineering simulation grids.

Cellular composition & symmetry

Mixed geometry, lighting, and math. SCE Cell Sym: 109.1 ms → 1.0 ms (109× faster, 99.7% energy reduction). Composite scene structure.

Arithmetic-intensive compute

Mandelbrot kernel: 146 ms → 3 ms (49×) on an Nvidia GPU on AWS. GPU utilization fell from 93% to 37%, the same hardware finishing the same work in a fraction of the duty cycle.

These workloads share no common source code, no common domain, and no common hardware target. What they share is a CJD architecture that expressed what each workload needed to accomplish, and wantware generated the execution that realized it.

20–114×
Workload acceleration range
90–99%
Energy reduction per workload
60+
Composite workloads optimized
0
Code changes required
On result variance: CJDs guarantee systematic, real-time exploration of the viable execution space on the target hardware, not a specific multiplier. Results depend on memory-bandwidth profile, compute-to-memory balance, control-flow irregularity, and tensor topology. The ranges above are measured, not projected. Your numbers come from your workload on your hardware.

04 · The Structural Distinction

Why CJDs cannot be
evaluated like benchmarks.

Evaluating a CJD with benchmark methodology misunderstands what a CJD does. A benchmark produces a score. A CJD produces governance: an ongoing, adaptive execution policy that changes as hardware behavior changes. The relevant question is not "what score did it get?" but "what did the hardware deliver on this workload class under this governance?"

Dimension Traditional benchmark Composite Job Design
Unit of output A score or time measurement for a specific workload Adaptive execution policy across a class of work
Scope One workload type, fixed configuration Any workload class computing can perform
Hardware dependency Results are valid only for the hardware tested Same CJD adapts across NVIDIA, AMD, Intel via SPIR-V
Temporal validity Point-in-time; re-run required after any change Continuous; adapts without recompilation
Optimization timing Pre-tuned before measurement; frozen at deployment Synthesized at runtime against observed hardware behavior
What it tells you How fast one executable ran once on this hardware What this hardware can deliver on this work, continuously
Framework dependency Assumes CUDA, ROCm, or vendor SDK No CUDA. No ROCm. No framework dependency.

This is not a claim that CJDs are better than benchmarks at the benchmark's own job. CJDs do not replace benchmarks for hardware comparison or procurement decisions. The claim is narrower and more consequential: when the goal is production efficiency rather than hardware selection, benchmarks are the wrong tool. CJDs are the right one.


05 · Pilot Process

How CJDs are generated,
validated, and deployed.

Chameleon® pilots produce CJDs for operator-nominated workloads and validate the resulting execution against the operator's own hardware. The process is structured to deliver first evidence within 30 days and full pilot artifacts within 90 days.

01
Workload nomination

The operator nominates representative workloads: AI inference, training, simulation, media processing, or scientific compute. MindAptiv scopes the CJD against the declared hardware target, OS, and objective (performance, energy, throughput, or latency).

02
CJD generation

Chameleon® analyzes the workload input (natural language description, GLSL kernels, or a workload harness) and generates a CJD that describes the work classes in scope. No refactoring. No framework lock-in.

03
Instruction synthesis

Wantware synthesizes optimized SPIR-V from the CJD against the target hardware. Kernel fusion, scheduling rebalancing, and memory-access reshaping are applied continuously, not once at compile time.

04
Measurement and artifacts

Results are measured against the operator's baseline: workload throughput, GPU power draw, GPU utilization, and wall-clock time. Pilot artifacts include a telemetry snapshot, method notes, and a repro guide for independent verification.

05
Production pathway

Validated CJDs become the execution substrate for the workload class. Integration is drop-in: wantware operates beneath the framework layer, preserving existing application investments.

Pilot engagements run on cloud, on-premises, or via hosted access. AWS and OCI GPU instances are the recommended starting point; on-prem runs on any NVIDIA or AMD hardware with Vulkan support. Setup time is 15–25 minutes on first run.


06 · Strategic Implications

The shift from benchmark performance
to execution governance.

Benchmark-driven infrastructure procurement has produced a generation of datacenters that are measurably fast on their benchmark suite and measurably inefficient on their production workloads. GPU clusters running at 10–35% effective utilization are the direct result of optimizing for benchmark performance rather than production execution.

CJDs represent a reframing of the infrastructure efficiency problem: from "which hardware scores highest on this benchmark" to "how do we extract maximum useful work from the hardware we have?" The financial implication is direct. A hyperscaler running 10,000 GPUs at 45% utilization carries the hidden debt of 5,500 purchased-but-idle GPUs, roughly $220M in latent capital at current pricing. Moving toward 90%+ utilization via CJD-governed execution means the same physical fleet delivers materially more output, deferring or eliminating the next procurement cycle.

The compounding effect matters. A 60× acceleration on a training workload does not only save time; it reduces GPU-hours required, which reduces energy draw, which reduces cooling load, which reduces facility cost, which defers capacity expansion. Execution efficiency improvement is a multiplier on capital deployed, not a linear cost reduction.

"The differentiated value is not in the silicon; it is in how efficiently the silicon is used. Execution-layer control is the structural moat."

Silicon procurement cycles are also affected. Each hardware generation transition (across Nvidia GPU generations, or from Nvidia to AMD to custom ASICs) requires re-optimization of conventional software stacks. For large operators, that is a 6–18 month engineering investment per transition per workload class. CJDs governed by wantware adapt at the execution layer automatically. The CJD does not change when the hardware changes.

See what CJDs deliver on your hardware.

Tell us your target hardware, OS, and objective. MindAptiv confirms the fastest pilot path, usually within a day.