AIec

Benchmarks

Every number on this page is published with the artifact it came from, and one optimization is published because measurement rejected it. Where a figure exists only in a hand-written report, it is marked as weaker evidence rather than presented with the same confidence as a harness run.

Method

Baseline and after runs are matched on a conditions fingerprint — the same hardware, the same workers, the same images, the same runtime. That fingerprint is 602f511b90348c50. Twenty-five attempts were made per side.

Baseline side

22 of 25 attempts succeeded. The three failures were stale-generation placement errors, and they are part of the result rather than a discard.

After side

25 of 25 attempts succeeded, with none of the baseline's stale-generation failures.

One statistic convention

The harness reports both p50_s and median_s for the same runs, and the hand-written reports mix the two. This page uses p50 throughout, read straight from the artifacts. A table transcribed from prose can differ from this one in the second decimal; that is why the difference is recorded here rather than silently averaged away.

Measured results

Matched before and after: 25 samples per side, same host, workers, images and fingerprint.
MetricBeforeAfterChangeSamples b / aNotesEvidence
Run, end to end
s
25.5014.56-42.9%22 / 25Submit to settled record. Matched conditions fingerprint on both sides.machine-checked artifact
Run, 95th percentile
s
28.1116.12-42.6%22 / 25The tail, which is where a cleanup bug shows up.machine-checked artifact
Run wall clock, client-observed
s
23.7712.71-46.5%22 / 25The same runs, measured from the client.machine-checked artifact
Cleanup phase
s
13.353.13-76.5%22 / 25The phase that was three quarters of the tail.machine-checked artifact
Placement phase
s
5.104.70-7.9%22 / 25Reservation plus the machine actually being built.machine-checked artifact
Task phase
s
0.870.70-19.0%22 / 25The workload itself. Included as a control: it should not change.machine-checked artifact
Queue wait
s
8.227.63-7.2%22 / 25Submission to a worker accepting the work.machine-checked artifact
Client wait no phase claims
s
5.995.96-0.5%22 / 25The honesty check on the rows above: if the gain had been moved into the unaccounted bucket, this row would have grown.machine-checked artifact
Control-plane requests per Run
requests
3122-29.5%22 / 25Accounting delta divided by runs. A cost that should fall when cleanup is fixed.machine-checked artifact
Sequential soak
runs
————100 / 100 succeededmachine-checked artifact
Parallel soak
runs
————77 / 80 succeededmachine-checked artifact
Guard gateway CPU, per 64 model requests
s
————0.46 s of gateway CPU, against 0.00 s for the same traffic unguardedmachine-checked artifact
Guard gateway memory, above idle
KiB
————+988 KiB above idle, +1 804 KiB above startupmachine-checked artifact

Change is before to after. Machine-checked artifact means the value is readable from a named key in a committed harness report; prose only means it exists solely in a hand-written report.

What the numbers qualify

Run, end to end
The three baseline failures were stale-generation placement errors, so the before side is 22 samples of 25 attempts. The after side succeeded 25/25.
Cleanup phase
Placement moved only −7.9 % over the same run, so this is a cleanup result and not a placement result.
Placement phase
Decomposed on the after side as scheduler 1.946, allocation 0.515, boot 0.534, workspace 0.000. `workspace` is 0 because this workload clones no repository; on a Run that does it is 1.2-1.3 s, which is what the rejected repository cache was aimed at. `placement.*` names are dotted sub-phases of `phase_placement`, so any tool that sums phases must exclude them.
Task phase
The command is `true`, so this is process spawn and teardown rather than real work. It is not a workload throughput claim.
Client wait no phase claims
Flat, which is the point: the improvement is inside accounted phases.
Control-plane requests per Run
Compared like every other row: the baseline file against the matched after file, both under conditions fingerprint 602f511b90348c50. The block counts HTTP requests to the control plane, not database queries, so it is a transport cost and not a query cost.
Sequential soak
run_cleanup_failures empty; available vCPUs 12.0 → 12.0; non-terminal sandboxes at the end 0.
Parallel soak
The three failures are `backend unavailable: transient: no schedulable worker has capacity` at four-way concurrency. Their cause is unattributed and the soak has not been rerun since. The endpoint census returns to baseline, so this is not an accumulation, but the intermediate checkpoints do not: the harness observed up to 3 non-terminal sandboxes and down to 9.0 available vCPUs mid-soak. This is the least flattering number we publish and it is the one on the front page.
Guard gateway CPU, per 64 model requests
64 requests per sample, concurrency 1, 1 MiB responses, 16 384-byte provider chunks, 2 warmup requests per side, a fresh gateway per pair and alternating order. Local HTTP in a disposable network namespace: not WAN, not TLS, and excluding guest, nftables, watcher and control-plane server. Baseline wall time 0.176 s against 1.614 s guarded. Conventional median only; no confidence interval and no percentile claim.
Guard gateway memory, above idle
RSS sampled every 5 ms, with VmHWM recorded alongside it.

Soaks

Sequential — 100 / 100

One hundred runs, one at a time. Every run succeeded and the resource census after the run matched the census before it: no leaked containers, no leftover TAP devices, no stranded directories, no non-terminal sandboxes. This is the measurement that says the system does not degrade under sustained load, and it is the one that was fixed by correcting a bookkeeping bug rather than by optimizing anything.

Parallel — 77 / 80

Twenty batches of four. Three runs were refused with no schedulable worker has capacity. This is the least flattering published number on this page and it is here for that reason: at four-way concurrency the admission gate refuses work when every worker is at its measured capacity, rather than oversubscribing the host. The census still balanced. Those three refusals have not been fully explained.

Guard overhead

Measured against an unguarded control across seven paired samples, on a local network connection. Median CPU and RSS for the governed and unguarded cases are in the table above.

These numbers are narrower than they look

Local HTTP, not a real network. The measurement excludes the guest, the nftables rules themselves, the watcher, and all control-plane cost. There is no confidence interval. Read this as "the policy evaluation itself is cheap on this hardware under these conditions", and nothing stronger.

The optimization we threw away

Repository object cache: off by default

Expected: less network traffic per Run, because a shared object store would replace a fresh shallow clone.

Measured: 26× more network bytes and roughly two seconds slower per Run. Guest receive bytes were 9 280 on first use and 9 654 median thereafter with the cache off, against 260 690 first and 253 223 median with it on.

Why: the cache re-resolved the repository ref on every Run, including a cache hit, to re-check that a repository which was public is still anonymously readable. In a calibrated guest, git clone --depth 1 moved 8 580 bytes while git ls-remote -- <url> HEAD moved 250 950. The safety check cost 29× the entire shallow clone it was protecting.

The guard was correct and the optimization was wrong. It ships off, and stays off until the re-check is redesigned. A cache bypass run matched the cache-off phase exactly, which is the authorization fix working.

Weaker evidence

These figures are published because they describe real work, but no JSON artifact backs them and they were recorded by hand.

Reproducing this

The harness scripts live in benchmarks/ and the operational validation scripts in scripts/. The committed JSON reports are the inputs this page reads.

python3 -m pytest benchmarks/tests/ -q     # benchmark harness tests
These claims cannot be regenerated by accident

web/data/benchmarks.json is reviewed by hand and is never auto-regenerated by a benchmark run. A new measurement does not silently rewrite the project's published numbers; someone has to read the artifact and decide whether the claim is still true.