Benchmarks
Every number on this page is published with the artifact it came from, and one optimization is published because measurement rejected it. Where a figure exists only in a hand-written report, it is marked as weaker evidence rather than presented with the same confidence as a harness run.
Method
Baseline and after runs are matched on a conditions fingerprint — the same
hardware, the same workers, the same images, the same runtime. That fingerprint
is 602f511b90348c50. Twenty-five attempts were made per side.
Baseline side
22 of 25 attempts succeeded. The three failures were stale-generation placement errors, and they are part of the result rather than a discard.
After side
25 of 25 attempts succeeded, with none of the baseline's stale-generation failures.
The harness reports both p50_s and median_s for
the same runs, and the hand-written reports mix the two. This page uses
p50 throughout, read straight from the artifacts. A table
transcribed from prose can differ from this one in the second decimal; that is
why the difference is recorded here rather than silently averaged away.
Measured results
| Metric | Before | After | Change | Samples b / a | Notes | Evidence |
|---|---|---|---|---|---|---|
| Run, end to end s | 25.50 | 14.56 | -42.9% | 22 / 25 | Submit to settled record. Matched conditions fingerprint on both sides. | machine-checked artifact |
| Run, 95th percentile s | 28.11 | 16.12 | -42.6% | 22 / 25 | The tail, which is where a cleanup bug shows up. | machine-checked artifact |
| Run wall clock, client-observed s | 23.77 | 12.71 | -46.5% | 22 / 25 | The same runs, measured from the client. | machine-checked artifact |
| Cleanup phase s | 13.35 | 3.13 | -76.5% | 22 / 25 | The phase that was three quarters of the tail. | machine-checked artifact |
| Placement phase s | 5.10 | 4.70 | -7.9% | 22 / 25 | Reservation plus the machine actually being built. | machine-checked artifact |
| Task phase s | 0.87 | 0.70 | -19.0% | 22 / 25 | The workload itself. Included as a control: it should not change. | machine-checked artifact |
| Queue wait s | 8.22 | 7.63 | -7.2% | 22 / 25 | Submission to a worker accepting the work. | machine-checked artifact |
| Client wait no phase claims s | 5.99 | 5.96 | -0.5% | 22 / 25 | The honesty check on the rows above: if the gain had been moved into the unaccounted bucket, this row would have grown. | machine-checked artifact |
| Control-plane requests per Run requests | 31 | 22 | -29.5% | 22 / 25 | Accounting delta divided by runs. A cost that should fall when cleanup is fixed. | machine-checked artifact |
| Sequential soak runs | — | — | — | — | 100 / 100 succeeded | machine-checked artifact |
| Parallel soak runs | — | — | — | — | 77 / 80 succeeded | machine-checked artifact |
| Guard gateway CPU, per 64 model requests s | — | — | — | — | 0.46 s of gateway CPU, against 0.00 s for the same traffic unguarded | machine-checked artifact |
| Guard gateway memory, above idle KiB | — | — | — | — | +988 KiB above idle, +1 804 KiB above startup | machine-checked artifact |
Change is before to after. Machine-checked artifact means the value is readable from a named key in a committed harness report; prose only means it exists solely in a hand-written report.
What the numbers qualify
- Run, end to end
- The three baseline failures were stale-generation placement errors, so the before side is 22 samples of 25 attempts. The after side succeeded 25/25.
- Cleanup phase
- Placement moved only −7.9 % over the same run, so this is a cleanup result and not a placement result.
- Placement phase
- Decomposed on the after side as scheduler 1.946, allocation 0.515, boot 0.534, workspace 0.000. `workspace` is 0 because this workload clones no repository; on a Run that does it is 1.2-1.3 s, which is what the rejected repository cache was aimed at. `placement.*` names are dotted sub-phases of `phase_placement`, so any tool that sums phases must exclude them.
- Task phase
- The command is `true`, so this is process spawn and teardown rather than real work. It is not a workload throughput claim.
- Client wait no phase claims
- Flat, which is the point: the improvement is inside accounted phases.
- Control-plane requests per Run
- Compared like every other row: the baseline file against the matched after file, both under conditions fingerprint 602f511b90348c50. The block counts HTTP requests to the control plane, not database queries, so it is a transport cost and not a query cost.
- Sequential soak
- run_cleanup_failures empty; available vCPUs 12.0 → 12.0; non-terminal sandboxes at the end 0.
- Parallel soak
- The three failures are `backend unavailable: transient: no schedulable worker has capacity` at four-way concurrency. Their cause is unattributed and the soak has not been rerun since. The endpoint census returns to baseline, so this is not an accumulation, but the intermediate checkpoints do not: the harness observed up to 3 non-terminal sandboxes and down to 9.0 available vCPUs mid-soak. This is the least flattering number we publish and it is the one on the front page.
- Guard gateway CPU, per 64 model requests
- 64 requests per sample, concurrency 1, 1 MiB responses, 16 384-byte provider chunks, 2 warmup requests per side, a fresh gateway per pair and alternating order. Local HTTP in a disposable network namespace: not WAN, not TLS, and excluding guest, nftables, watcher and control-plane server. Baseline wall time 0.176 s against 1.614 s guarded. Conventional median only; no confidence interval and no percentile claim.
- Guard gateway memory, above idle
- RSS sampled every 5 ms, with VmHWM recorded alongside it.
Soaks
Sequential — 100 / 100
One hundred runs, one at a time. Every run succeeded and the resource census after the run matched the census before it: no leaked containers, no leftover TAP devices, no stranded directories, no non-terminal sandboxes. This is the measurement that says the system does not degrade under sustained load, and it is the one that was fixed by correcting a bookkeeping bug rather than by optimizing anything.
Parallel — 77 / 80
Twenty batches of four. Three runs were refused with
no schedulable worker has capacity. This is the least flattering
published number on this page and it is here for that reason: at four-way
concurrency the admission gate refuses work when every worker is at its
measured capacity, rather than oversubscribing the host. The census still
balanced. Those three refusals have not been fully explained.
Guard overhead
Measured against an unguarded control across seven paired samples, on a local network connection. Median CPU and RSS for the governed and unguarded cases are in the table above.
Local HTTP, not a real network. The measurement excludes the guest, the nftables rules themselves, the watcher, and all control-plane cost. There is no confidence interval. Read this as "the policy evaluation itself is cheap on this hardware under these conditions", and nothing stronger.
The optimization we threw away
Expected: less network traffic per Run, because a shared object store would replace a fresh shallow clone.
Measured: 26× more network bytes and roughly two seconds slower per Run. Guest receive bytes were 9 280 on first use and 9 654 median thereafter with the cache off, against 260 690 first and 253 223 median with it on.
Why: the cache re-resolved the repository ref on every Run,
including a cache hit, to re-check that a repository which was public is still
anonymously readable. In a calibrated guest, git clone --depth 1
moved 8 580 bytes while git ls-remote -- <url> HEAD moved
250 950. The safety check cost 29× the entire shallow clone it was protecting.
The guard was correct and the optimization was wrong. It ships off, and stays off until the re-check is redesigned. A cache bypass run matched the cache-off phase exactly, which is the authorization fix working.
Weaker evidence
These figures are published because they describe real work, but no JSON artifact backs them and they were recorded by hand.
- Artifact collection. 129 360 ms with one read per 64 KiB chunk against 8 945 ms with bounded groups of 32, and 39 ownership calls for a 1 MiB collection because the chunk endpoint brackets every read with two lease checks. This is the mechanism behind the per-Run request-cost row.
- Repository cache calibration. The 8 580 and 250 950 byte counts are single-guest readings with the interface counter read either side of one command, not a distribution.
Reproducing this
The harness scripts live in benchmarks/ and the operational
validation scripts in scripts/. The committed JSON reports are the
inputs this page reads.
python3 -m pytest benchmarks/tests/ -q # benchmark harness tests
web/data/benchmarks.json is reviewed by hand and is never
auto-regenerated by a benchmark run. A new measurement does not silently
rewrite the project's published numbers; someone has to read the artifact and
decide whether the claim is still true.