What has been measured, and how to test it yourself.
A precise account of what Chameleon does, what was measured and under what conditions, why fixed benchmarks capture only part of an adaptive system, and a one-month evaluation that your own team runs on your own workloads.
Measured to date: 18 workloads on one AMD GPU, plus NVIDIA runs on three clouds.
This brief is for infrastructure and performance engineers who want to know what Chameleon does, what has been measured, and how to test it on their own workloads before changing anything in production. The Phase 1 pilots were not benchmarks. Their purpose was to show, independently and across different architectures, what Composite Job Designs make possible. In plain terms:
- The results agree across independent environments. A university team (DEHub, on-premises AMD integrated GPU), runs recorded by PREDICTif Solutions on AWS, and our own runs on Oracle Cloud and Google Cloud (NVIDIA) all show the same pattern: large speedups and large energy reductions against the first run of the same job, in overlapping ranges. The DEHub and AWS figures are not our numbers, and they track with the OCI and GCP figures. What carries the conclusion is that consistency across different architectures, not any single run or report.
- Measured by a university team. Rowan University’s Digital Engineering Hub (DEHub) ran 18 workloads on one integrated AMD GPU. We supplied the executable and had no access to the system. The resulting report is a draft and has not been signed. We proposed the Phase 1 scope, DEHub management agreed to it, and we delivered what was agreed. We said at the outset that two milestones, Buffer Feeds and Synergy, were unfunded at the time and outside that scope. Management then requested a Phase 2 proposal, reviewed it and accepted it, pending completion of those two milestones.
- Run on cloud NVIDIA GPUs. The AWS runs were independently recorded by PREDICTif Solutions (since acquired by nClouds), in an AWS-funded project. Amazon has approved funding for the current AWS phase (AWS Phase 2), which is active. We also report runs on Oracle Cloud and Google Cloud instances (T4, A10G, A10, A100). The Oracle and Google runs are our own results.
- Evaluation scope. The evaluation binaries are built for evaluation. Single-GPU runs are the measured scope; multi-GPU operation and long-running use are next steps. Buffer Feeds and Synergy were unfunded during Phase 1, so they are not part of what was measured.
The two milestones
Two milestones remain, and Phase 2 is pending their completion. Neither was funded during Phase 1, and neither is part of what was measured.
- Buffer Feeds (Milestone 01, Distribution). Distributes work across all components within a single node, including multi-GPU configurations such as the NVIDIA A100 and B100, with no hypervisor dependency at the node level. It enables single-node, multi-GPU Composite Job Designs. It does not by itself run work across multiple nodes: cross-node distribution is a separate networking step that follows. Every measurement in this brief is single GPU, so none of it is attributable to Buffer Feeds.
- Synergy Expansion (Milestone 02, Intelligence). Expands Essence’s capacity to understand, create, edit and materialize Aptivs across the full range of what computing can do, assimilating software’s historical workarounds as governed, direct capability. It includes natural-language creation of workloads, with generative AI proposals governed by Meaning Coordinates.
What is available today
| Capability | Status |
|---|---|
| Runtime generation of SPIR-V, executed through Vulkan | Available now in the pilot build |
| SPIR-V file export | Not available. Re-established during the first 30 days of funded milestone work |
| OS-hosted runtime, illumin8 and the other Aptivs | Not running today. They return when the milestones are complete |
| Existing CUDA or PyTorch workloads, as submitted | Not accepted by the current pilot build |
| Changes to the pilot build | Not before funding is in place to complete the two milestones |
| Fixed problem sets on your hardware | Available now in the pilot build |
| Your own workloads, compared with your production build | After the two milestones are complete |
| Single-GPU runs | The measured scope |
| Multi-GPU operation (Buffer Feeds) | Waits for Milestone 01 |
| Natural-language creation of workloads (Synergy Expansion) | Waits for Milestone 02 |
| Signed DEHub Phase 1 report | Draft, not yet signed. DEHub accepted the Phase 2 proposal, pending the two milestones |
The sections that follow set out what was measured and under what conditions, then propose a one-month evaluation that your own team runs and controls.
From intent to SPIR-V, executed through Vulkan.
Chameleon takes a statement of intent, expresses it as a Composite Job Design (CJD), and uses Morpheus® to generate SPIR-V instructions for the target GPU at runtime. It is not a compiler pass, a fixed kernel library, a prompt-to-code assistant or a profiler.
- Inputs. The current pilot build runs a fixed set of problem sets. Our pilot page describes the input forms the platform is designed to take (natural language, existing GLSL, kernels, or a workload harness with a stable execution boundary) and states no CUDA, ROCm, oneAPI or PyTorch dependency.
- Composite Job Design. A CJD describes a set of valid, semantically equivalent execution designs for one intent, bounded by declared constraints such as power or latency. Per the DEHub report, any adaptation selects among pre-validated paths; intent and constraints do not change at runtime.
- Two execution modes. Fixed: a single design is selected and used unchanged. Bounded adaptive: several realizations are evaluated under controlled conditions and converge to a stable design. Chameleon generates the SPIR-V at runtime in both modes. Exporting SPIR-V as a file is not available yet.
- What it restructures. Kernel boundaries, instruction ordering, parallelism granularity, memory access patterns, scheduling and synchronization.
- Footprint in the DEHub evaluation. One Linux executable of about 14 MB. No agents, background services or SDKs. Ubuntu 24.04, kernel 6.11 or newer (needed for that AMD platform), the open AMDGPU driver, Vulkan 1.3 or newer.
- Hardware path. SPIR-V through Vulkan targets NVIDIA, AMD and Intel GPUs. NVIDIA T4, A10G, A10, A100 and H100 runs are reported on cloud instances (H100 on Google Cloud only); the DEHub evaluation covers only an AMD Radeon 8060S.
Why SPIR-V and not CUDA, ROCm or oneAPI?
CUDA, ROCm and oneAPI each lead to one vendor's hardware, so work written for one has to be ported and tuned again for another. SPIR-V is the cross-vendor format that NVIDIA, AMD and Intel drivers accept, and each driver lowers it to the native instructions of its own chip. Because Chameleon generates the instructions instead of asking an engineer to write them, it can target that shared layer and tune the output for each machine. This path does not replace vendor libraries, and DirectX 12 and Metal use their own intermediate formats. The Why SPIR-V page explains the stack for engineers who work in CUDA, with diagrams.
Tuning that is usually done by hand
Chameleon uses Morpheus® to generate the machine instructions from the CJD. Performance engineers usually tune the tasks below by hand, workload by workload, and the result holds for one workload on one chip at one moment. When the hardware or the conditions change, the work starts again. Here the same tasks are generated for the hardware present.
Output carried as a bitstream
Output from a CJD is carried in the Wantverse [.wv] format, which the technical docs describe as a stream and not a file in the conventional sense. The bitstream is designed to be laid out contiguously so memory does not fragment, to be headerless so the payload is the whole stream, and to be encrypted per stream by StreamWeave®, which shifts algorithm combinations continuously. Jobs are also woven together along the stream with no fixed start or end point, which adds a form of obfuscation. The job details needed for processing are carried by Meaning Coordinates inside the Aptivs, not by a header. File-based output for pipelines that need a file is not running today. It is re-established during the first 30 days of funded milestone work. The docs list StreamWeave per-stream encryption as dependent on the Synergy milestone.
Essence adapts. A benchmark assumes the opposite.
A standard benchmark assumes a fixed tool and a fixed hardware environment: a defined program, run on a named machine, scored against another program on the same machine. That design suits the first two eras of software. It does not suit a system whose purpose is to adapt.
| Era 1 · code-driven | Era 2 · AI generation | Era 3 · intent-native (Essence) | |
|---|---|---|---|
| What is produced | One program per target, written by people. | Code proposed by a model. It varies from run to run and is a fixed artifact once accepted. | Execution resolved from intent when it runs. There is no single artifact per target. |
| Route to the result | One, chosen by the developer. | Probabilistic: many candidate routes. | Deterministic, resolved against the priorities the user declares. |
| Fit to the hardware present | Ports, frameworks and hypervisors bridge code written for one machine shape to the machine actually present. | Adaptation means generating new code for the new target. | Resolved from the intent at execution, for the hardware actually present. |
| What a benchmark captures | The speed of the artifact on that machine. | The same, one output at a time. | One machine’s resolution of the intent. Not the capability. |
What this means for reading the results
- Demonstration points, not scores. The DEHub, AWS, Google Cloud and OCI figures show Essence adapting on specific workloads and specific hardware. They are not benchmark scores, and no fixed-workload ranking is claimed.
- Same intent, different execution, by design. On other hardware, or with other declared priorities, Essence resolves the same intent differently. Two runs on two machines compare two resolutions, not one program twice.
- The first-run baseline follows from this. The speedup compares Essence’s first resolution of a job (Composite Job Design 1) with its adapted result, because there is no fixed program to hold constant.
- A fixed comparison is still a useful reference. Buyers need one, so the evaluation below includes a comparison against your own frozen build once your own workloads can run. It measures one slice. The questions that fit better are whether the outcome meets the declared priority on the hardware present, and whether it holds when the hardware or conditions change, without rewriting.
Why Phase 2 uses the evaluator’s own workloads
Because a benchmark scores a fixed program, a benchmark suite cannot show what a Composite Job Design does. DEHub’s Phase 2 therefore takes a different form. It is not a request for a score, and it does not signal doubt about CJDs. DEHub wants to run its own workloads as CJDs, so that the result answers the question the institution cares about: what happens on its own work, on its own hardware, against success criteria it fixes in advance. We expect the same to hold for other evaluators. A fixed comparison against a frozen build becomes possible once your own workloads can run. Until the milestones are complete, the current build runs a fixed set of problem sets. The CJD run on your own workload is the measurement that matters.
Seventeen of 18 workloads ran 22× to 114× faster than their first run. One ran 14× faster.
Read these as demonstration points rather than benchmark scores (see the previous section). The table lists each result with its source and hardware. In every case, speedup is measured against the first run of the same job.
| Result | Source and hardware | Who measured | Notes |
|---|---|---|---|
| 14.0× to 114.1× across 18 workloads (median 39.7×). Energy reduction 97.8% to 99.7% (median 99.2%). Figures from Table 5 of the report. | DEHub personnel, on-prem, AMD Radeon 8060S integrated GPU, 1920 × 1080, Ubuntu 24.04. | Measured by third-party staff on their own system. We had no access. | Report is in draft. See the check below. |
| Cloud NVIDIA GPUs: 3.9× to 53.6× speedup across 7 cloud and GPU combinations, median about 28× to 32× on each. Energy reduction 75.3% to 99.6%. Per-platform figures are in the table below. | Our pilot log (revision 11): AWS, Google Cloud and OCI, same executable. | AWS runs: recorded independently by PREDICTif Solutions (since acquired by nClouds), in an AWS-funded project (AWS Phase 2 funding approved by Amazon, project active). Google Cloud and OCI runs: our own results. | Power draw on the T4 rows barely changes (for example 63 W to 60 W). The energy saving comes from shorter runtime. Variance between clouds is attributed to host differences. Speedup compares the best of runs 2 to 6 with the first run of the same job. |
| 103× on a 3D rendering job, 32 minutes to 18.8 seconds. | Our internal measurement, 2011 Mac Pro, CPU multithreading. | Internal. | An earlier CPU result that illustrates the same principle. It is not a GPU measurement. |
Results by cloud and GPU
Each row summarizes the tests in our pilot log (revision 11) for one GPU on one cloud. A test is one workload at one resolution. Speedup is the best of runs 2 to 6 against run 1 of the same job, shown as a range across tests with the median in brackets. Resolution is 3840 × 2160 unless noted; some AWS T4 tests use 1920 × 1080, and the H100 tests span several resolutions.
| Cloud | NVIDIA GPU | Tests | Speedup | Energy reduction | Recorded by |
|---|---|---|---|---|---|
| AWS | Tesla T4 | 24 | 10.8× to 53.6× (median 31.5×) | 93.0% to 98.9% | PREDICTif (independent) |
| AWS | A10G | 18 | 14.1× to 49.2× (median 32.4×) | 95.3% to 99.6% | PREDICTif (independent) |
| AWS | A100 40 GB | 6 | 15.0× to 31.4× (median 31.1×) | 95.6% to 98.0% | PREDICTif (independent) |
| Google Cloud | Tesla T4 | 21 | 10.0× to 49.4× (median 27.8×) | 90.3% to 98.7% | MindAptiv |
| Google Cloud | A100 40 GB | 20 | 9.3× to 32.8× (median 31.0×) | 94.2% to 98.8% | MindAptiv |
| Google Cloud | H100 80 GB | 16 | 14.0× to 34.0× (median 31.1×) | 77.3% to 98.2% | MindAptiv |
| OCI | NVIDIA A10 | 23 | 3.9× to 41.4× (median 32.2×) | 75.3% to 98.0% | MindAptiv |
- H100 results are from Google Cloud only. They span four resolutions, 3840 × 2160 to 7680 × 4320, over four workloads. Red Nebula holds 30.7× to 31.4× at all four resolutions. Firework Pop holds 32.3× to 32.4× through 6016 × 3384 and falls to 15.6× at 7680 × 4320, with energy reduction of 77%. Smoke_Riser ranges from 31.4× down to 14.0× (energy reduction 91% at the low end), and the SDF workload from 34.0× to 18.4×.
- No speedup is reported at 15360 × 8640. The four H100 tests at that resolution record the first run only.
- B100. No B100 run is recorded. The H100 setup carries over only in part: NVIDIA documents that Blackwell GPUs require the open GPU kernel modules, and the driver branch, the Vulkan driver package and the Vulkan version exposed must be confirmed on the B100 nodes against a working H100 node. A B100 figure itself would come from the evaluation. An A5000 and an AWS-hosted H100 are also not yet recorded.
- Not tabulated. Two single Google Cloud tests (an A10 and an A100 with no variant stated), because one test does not make a range.
- What the log measures. Time, power and energy per run. It has no output-comparison column. Where the log notes mention fidelity or quality, that is a description, not a measurement.
The rows agree with each other
The DEHub rows (14× to 114× speedup, 97.8% to 99.7% energy reduction), the cloud rows (3.9× to 53.6× speedup with a median near 30× on every platform, and 75.3% to 99.6% energy reduction) and the earlier CPU result all show the same direction, with overlapping magnitudes. They come from different people and different architectures (AMD integrated, and NVIDIA T4, A10G, A10, A100 and H100 cloud GPUs), and the DEHub and AWS measurements were not made by us. The purpose of the pilots was never a score. It was to show, across architectures, that the same Composite Job Design approach engineers a solution for each. Read together, the rows show that. No single report has to carry the conclusion.
Our bandwidth (WarpSpeed) and video (illumin8) results concern signal processing, not GPU execution, and are outside this brief. Our illumin8 technology forms a unified pipeline for signal processing, which enables adaptive quality, without loss. It derives detail mathematically from the signal, based on Meaning Coordinates, and the original can be reconstructed when required. illumin8 is not running as a build today. It returns when the milestones are complete, so it is outside the first evaluation. The pilot log behind this section records speed, power and energy only, with no output-comparison column. A quality measure for your workloads would therefore be agreed with you in advance and run once illumin8 is deployed again, and a reconstruction check (rebuild the original from the output and compare it with the input) can be included then.
Check run on the Phase 1 table
Each energy figure equals 1 minus (optimized power × optimized time) over (baseline power × baseline time), for all 18 rows. The energy column is therefore calculated from the recorded time and power readings, not read from a separate energy meter.
Where the measurements apply, and where your evaluation extends them.
Each item below defines the conditions the results were measured under. Your evaluation extends them to your own system.
- Workload type. The DEHub workloads are shader-style kernels run at 1920 × 1080, chosen as representative of data-pipeline patterns. The report relates them to simulation, imaging, robotics and AI tensor work. AI inference and training workloads are not yet part of the measured set.
- Baseline. The baseline is the first run of the same job. In our pilot log, run 1 of each test is Composite Job Design 1, described as the default implementation with the fewest permutations, and runs 2 to 6 are successive refinements. A speedup therefore measures how far Essence’s own refinement moves from its first output. Comparison with your existing code, a hand-tuned shader or a vendor compiler path is what the client-run evaluation below adds. The DEHub evaluation uses the same first-run baseline.
- Hardware. The headline 114× comes from a 40-compute-unit integrated GPU with shared memory. Discrete datacenter GPUs were measured on the cloud NVIDIA runs. No B100, A5000 or AWS-hosted H100 run is recorded yet. The H100 setup carries over to B100 only in part, because Blackwell needs the open kernel modules and a driver, Vulkan package and Vulkan version that are confirmed at scoping.
- Repeatability. The report states results were stable across runs. Our pilot log records one pass per test, up to six successive variants. The evaluation below specifies repeated runs with spread reported.
- Independence. The tests were run by DEHub personnel on their own system, and we had no access to it. We proposed the Phase 1 scope and DEHub management agreed to it. The Phase 2 proposal was prepared at management’s request, then reviewed and accepted by management, pending the two milestones. The results do not depend on that report alone: the AWS runs were recorded by a third party, and the OCI and GCP runs fall in the same range.
- Edge and physical AI. Edge computing and physical AI, such as robots and vehicles, run on many small machines with tight limits, and the company sees both as massive opportunities for businesses. Results on edge hardware have not yet been measured.
- Scope. Single GPU. Multi-GPU and cluster operation follow later milestones. Our guarantee is not a particular number. It is real time optimization of data ordering, scheduling, memory, networking and hardware parallelization for the workload and hardware at hand, which is what highly skilled engineers do by hand and cannot do in real time. Composite Job Designs are how it is done. Results depend on memory bandwidth, compute-to-memory balance, control-flow irregularity and tensor topology.
- Cloud runs. The AWS runs were recorded independently by PREDICTif Solutions, in a project funded by Amazon. The Google Cloud (including all H100) and OCI runs are our own results.
Adoption is staged, and every stage after the first evaluation depends on the funded milestones.
The documented path expands only where a workload justifies it. Teams can adopt per workload, per product or per team.
- Export Mode. The design lets Essence hand standard artifacts to an existing pipeline unchanged: Git, Jenkins, GitHub Actions, GitLab CI, Azure DevOps, SAST and DAST scanners, artifact registries and current deploy targets. Essence is not a code generator, and these artifacts are an interface to existing tooling, not its purpose. File-based output is not running today. It is re-established during the first 30 days of funded milestone work.
- Hybrid. Fixed, scan-ready builds are kept for compliance, while controlled runtime behavior is introduced for chosen workloads. Policy gates decide which workloads are fixed and which are adaptive.
- Essence Runtime Mode. Execution happens dynamically within Essence, with policy, traceability and audit controls. Fixed builds can still be exported when certification or offline deployment requires them, once file-based output is re-established. The OS-hosted runtime returns when the milestones are complete.
These phases describe the platform’s design. Today only the GPU optimization build runs, and the measurements in this brief come from it on a single GPU, so the evaluation below comes before any rollout planning.
One month, your hardware, your criteria. Your own workloads follow the milestones.
Your team operates the evaluation and sets the success criteria in advance. The current pilot build runs a fixed set of problem sets, so an evaluation today is run on those, on your hardware. Running your own workloads and comparing them with your frozen production build becomes possible once the two milestones are complete. Both stages follow the same method. The structure follows the Phase 2 proposal prepared at DEHub management’s request (about one month, single GPU, performance and energy analysis). Run counts and tolerances are proposals for your engineers to adjust. If your question is whether efficiency can be gained without regression on a live system, the form that answers it is a live side-by-side (shadow) evaluation of your full system, described below. It depends on the milestones, so it cannot be run on the current build.
- Scope (week 0). You supply the GPU model, the driver version and kernel module type, the Vulkan version reported by vulkaninfo, the OS and version, and the objective (performance, energy, throughput or latency). Today the workloads are our fixed problem sets. After the milestones, you add one to three of your own.
- Confirm what can run. We state in writing which problem sets the current pilot build runs. The build is not modifiable for individual evaluations, so an existing CUDA kernel or PyTorch workload is not accepted as submitted. Changes to the build depend on funding to complete the two milestones.
- Baseline. Today the baseline is the first run of the same job, the default design. It shows how far adaptation moves on your hardware. A comparison with your production build is not part of the current build’s scope. Once your own workloads can run, you freeze and record your current production build, compiler, flags and driver before any Chameleon run, and both sides run identical inputs.
- Metrics. Wall time per unit of work, throughput and GPU utilization. Power and energy come from your own telemetry (vendor tools or an external meter) and are reported as measured, not calculated from time and nominal power. Record the time and energy of every refinement run as well as the final one, so that the cost of getting to the final design counts against the final figure.
- Priority shift. Run the same intent with the declared objective changed (for example from time to energy) and record how the execution and the metrics change. Confirmed at scoping.
- Hardware change. If your fleet mixes GPU models, run the same package unmodified on each model and report the metrics per model as well as in aggregate. Confirmed at scoping.
- Runs. Fixed warm-up, then repeated runs. Proposal: at least 30 runs per configuration, with baseline and Chameleon interleaved once a production baseline is in play. Report the median and spread and keep every raw value.
- Correctness. Outputs are compared with the reference within tolerances you set before testing. The tolerance can be per output or in aggregate across the run (for example, an equivalence threshold of 99.9% judged over the whole workload), and you choose the metric. Our pilot log contains no output-comparison measurement, so this evaluation is where output equivalence is first measured. A tolerance failure counts against the result.
- Live side-by-side (after the milestones). Your production system keeps serving unchanged. The same live inputs also run through the Chameleon path in shadow, and its outputs are not served. Outputs of the two paths are compared in aggregate against your tolerance, with time, throughput and energy recorded on both. Results are reported per GPU model and overall. Because the full system is the unit under test, this needs your own workloads to run, which depends on the two milestones. No manual-coding bridge is offered to get there sooner.
- Isolation. You operate the system. We provide the executable and guidance only, with no access to the environment or data, as in the DEHub evaluation.
- Pre-set criteria. You write the success threshold and stop conditions before the first run.
- Deliverables. You keep the raw CSV, the configuration record and a short written report. Sharing any of it with us is by agreement.
Suggested rhythm: setup in week 1, runs in weeks 2 and 3, analysis in week 4. Multi-GPU operation (Buffer Feeds) and natural-language creation of workloads (Synergy) depend on two milestones that are not complete, so they are outside this first evaluation. So are the .wv stream and StreamWeave encryption, which are not part of a single-GPU kernel test and depend on an incomplete milestone.
Current answers, and what the evaluation settles.
| Question | Current answer | Status |
|---|---|---|
| What is the baseline for the speedups? | The first run of the same job: Composite Job Design 1, the default implementation with the fewest permutations. Your own baseline is set in the evaluation. | Answered; your baseline is used in the evaluation |
| Why not just run a standard benchmark? | Standard benchmarks score a fixed artifact on a fixed machine. Essence resolves execution from intent for the machine present and the priorities declared, so a fixed score captures one resolution only. The evaluation keeps a frozen-baseline comparison as one reference and adds priority-shift and hardware-change runs. | Addressed in the evaluation |
| Does it work with existing CUDA code or PyTorch? | Not as submitted. Documented inputs are GLSL, kernels or a workload harness, with no CUDA dependency. The current pilot build runs a fixed set of problem sets and is not modifiable, so an existing CUDA kernel or PyTorch workload is not accepted as submitted. Changes to the build depend on funding to complete the two milestones. | Not in the current build |
| Does it help AI training or inference? | Phase 1 covered shader-style workloads. AI workloads are a candidate for the client-run evaluation. | Evaluation candidate |
| Multi-GPU and clusters? | Buffer Feeds (single-node multi-GPU) is a later milestone. | Planned |
| Is execution deterministic and auditable? | Fixed mode uses a single design. Phase 1 executed deterministic paths. SPIR-V file export, which would allow external review and hashing, is not available yet. | Partly available; export pending |
| What is installed, and what does it touch? | One executable of about 14 MB, with no agents or services, and no change to the host’s security policy (per the report). | Documented |
| What are the runtime overhead and generation latency? | Measured in the client-run evaluation. | Evaluation metric |
| Does output quality hold, with no regression? | Not measured yet. The pilot log records time, power and energy and has no output-comparison measure. Equivalence is measured in the evaluation, against a metric and a tolerance you fix in advance, judged per output or in aggregate. Our illumin8 pipeline for signal processing is described as enabling adaptive quality without loss, and that statement is outside the GPU evidence in this brief. illumin8 is not running as a build today. | Not yet measured |
| Which GPUs have been measured? | AMD Radeon 8060S (DEHub); NVIDIA T4, A10G and A100 on AWS; T4, A100 and H100 on Google Cloud; A10 on OCI. No B100, A5000 or AWS-hosted H100 run is recorded yet. The H100 setup carries over to B100 only in part, because Blackwell needs the open kernel modules and a driver, Vulkan package and Vulkan version that are confirmed at scoping. | Partly covered |
| Where does it not help? | Gains depend on memory bandwidth, compute-to-memory balance, control-flow irregularity and tensor topology. The lowest DEHub result was 14×. In the cloud log the lowest was 3.9× (OCI A10, Red Nebula). | Documented |
| Is there signed third-party validation? | Independent results agree: DEHub, AWS (recorded by PREDICTif), OCI and GCP fall in overlapping ranges, and the DEHub and AWS numbers are not ours. DEHub: Phase 1 delivered what was agreed and its report is in draft. Management requested a Phase 2 proposal, reviewed it and accepted it, pending completion of two milestones. AWS runs: recorded independently by PREDICTif Solutions. | Consistent across four environments; report in draft |
| What do the cost-savings claims rest on? | Our ROI calculators are models built on the measured speedups. | Projections |
| Licensing, support and data handling? | Agreed at scoping. | At scoping |
Chameleon’s measured gains are large: speedups over the first run of the same job, across shader-style workloads on a single GPU, and they are consistent across a university system, AWS and two other clouds. The guarantee is real time optimization of data ordering, scheduling, memory, networking and hardware parallelization, not a particular number. Because Essence adapts to the hardware and priorities in front of it, a fixed benchmark captures only one resolution of an intent. The right next step is a bounded evaluation on your own hardware, and on your own workloads against a baseline you define once the milestones are complete.