REST API
Base path /. Except health/metrics, send Authorization:. JSON uses snake_case; errors are {"error":. Every response includes x-operation-id; clients may supply a UUID in that header, otherwise the server generates one. This correlation header is not yet propagated into structured logs.
Sandboxes#
POST /— create. Body:v1/ sandboxes image,cpu,memory_mb,disk_mb,timeout_seconds,network.enabled.environment.workspaceaccepts{"type":or"empty"} {"type":. Repository credentials in URLs are rejected."git", "repo": "https://...", "reference": "main", "shallow": true} environment.toolkitsis a bounded list of namedsetup_commandsargv arrays.environment.guardselects an out-of-guest policy, with the same shape asresources.guardbelow. It replacesnetwork.enabledrather than accompanying it. A Firecracker request with noenvironment.guardgets the default no-network policy; a request that asks fornetwork.enabledwithout selecting a policy is refused unless the operator has setAIEC_ALLOW_LEGACY_NETWORK=1. Model credentials are configured on the worker and never travel in this document; the guest receivesplaceholder://<binding>instead.GET /— one bounded page of the tenant's sandbox history, newest first:v1/ sandboxes limit(default 50, store-clamped to 200) and the optionalafter_created_at+after_idcursor, which must be given together. It returns{sandboxes,, orderednext} created_at DESC,;id DESC nextis the cursor for the following page and isnullexactly on the last one, so a page says whether it truncated rather than looking complete. A destroyed sandbox stays in the history — teardown is a state transition, not a delete — so this is the tenant's whole history and it grows with tenure. Exposed asclient.list_all_sandboxes(...)andclient.list_sandboxes_after(...)in the Rust client, and assandboxes.list_all()andsandboxes.list_after(...)in the Python SDK.GET /— tenant-scoped detail.v1/ sandboxes/ {id} DELETE /— destroy.v1/ sandboxes/ {id} POST /— explicit lifecycle operations.v1/ sandboxes/ {id}/ start| pause| stop| resume pauseis supported by Firecracker PATCH semantics; development Bubblewrap returns501. Rust SDK:AIecClient::(cratepause(id) aiec-client) /resume(id); Python SDK:Sandbox.pause()/Sandbox.resume().POST /— return boundedv1/ sandboxes/ {id}/ git/ diff git diff --binaryfor a running Git workspace.POST /— bodyv1/ sandboxes/ {id}/ exec commandargv,working_directory,environment,timeout_seconds,stdin. Returns exit code, bounded stdout/stderr, duration, timeout flag.PUT /— bodyv1/ sandboxes/ {id}/ files path,content_base64, optionalmode.PUT|andDELETE / v1/ sandboxes/ {id}/ secrets/ {name} GET /manage one-hour, per-sandbox process environment secrets. Values are never returned by metadata endpoints.v1/ sandboxes/ {id}/ secrets GET /— list.v1/ sandboxes/ {id}/ files?path=/ workspace GET /— download.v1/ sandboxes/ {id}/ files/ content?path=/ workspace/ x POST /— create directory.v1/ sandboxes/ {id}/ files/ mkdir DELETE /with JSONv1/ sandboxes/ {id}/ files {path}— delete file.POST /,v1/ sandboxes/ {id}/ snapshots GET /.v1/ sandboxes/ {id}/ snapshots POST /— creates a new sandbox and returns it.v1/ snapshots/ {id}/ restore DELETE /.v1/ snapshots/ {id} POST|provides tenant-scoped upload/download/delete with a 64 MiB decoded-size limit.GET| DELETE / v1/ sandboxes/ {id}/ artifacts/ {name} GET /lists sorted object metadata for the development filesystem backend. The production S3 backend returns a typedv1/ sandboxes/ {id}/ artifacts 501for listing until ListObjectsV2 support is implemented and verified.The listing is bounded by
MAX_LISTED_OBJECTS(1000) objects under the sandbox's prefix, and is refused with413 limit_exceededpast that rather than truncated — a shorter list would be indistinguishable from a complete one. The bound is on the walk rather than on uploads:PUTof distinct names is uncapped, so a sandbox can hold more than one listing will describe. On the filesystem backend the listing also computes each object's size and SHA-256 by reading it, so its cost is the bytes of the prefix, not just its directory entries; the bound exists to keep that proportional to the platform rather than to tenant storage.GET /— metric totals for current tenant.v1/ usage
The Python SDK exposes the same size-guarded partial path as Sandbox.upload_artifact, download_artifact, and delete_artifact; listing remains unsupported, and the production S3 path remains unverified.
Runs#
POST /— create and drive a run to a terminal state, returning the settled run. Body:v1/ runs workload(image,command,setup,validations,artifacts,environment,secrets,timeout_seconds,git_evidence,repo),resources,requirements,retention,max_attempts,idempotency_key,requested_runtime,retained_seconds. The handler is synchronous on purpose: it places a machine, runs the work, collects artifacts and reclaims the machine before answering, so the response is the record of what happened.High-risk tool calls are approved in two phases by different principals.
POST /(scopev1/ sandboxes/ {id}/ guard/ approval SandboxesWrite) asks about one call and always answers immediately withapproved:when nothing has decided it, recording the request asfalse pending.GETandPOST /(scopev1/ sandboxes/ {id}/ guard/ tool-approvals GuardApprove) are the operator's queue and the decision itself. The ask carriesdigest, a SHA-256 over the canonical JSON of the call's arguments; it is required, because an approval that is not bound to one call is a capability rather than an approval. A grant is spendable once, only for the digest it was made for, and only by the key that asked. Self-approval is refused however much authority the key holds. See GUARD.md.
Both Guard history listings are paged the same way, for the same reason. A
proposal row is never reclaimed, and a decided tool approval is deliberately
retained, so a sandbox left running grows both without bound. GET / (scope GuardRead) and GET / (scope GuardApprove) take the same
limit (default 50, store-clamped to 200) and the same paired
after_created_at + after_id cursor, and return {proposals, and
{approvals, respectively, ordered created_at DESC,, with
next non-null exactly when another page follows. The id half of the cursor is
not decoration: created_at is not unique, and a keyset that cannot break a
tie either re-reads or skips the rows sharing a timestamp.
resources.guard selects an out-of-guest policy: either
{"policy_template":
with optional model_endpoint and allowlist inputs, or a full
{"policy": document. It replaces resources.network rather than
accompanying it — two descriptions of one egress would be two paths to the
internet — and it is enforced by the worker's Guard gateway, not by the
guest. A guarded run is placed on a microVM runtime; there is no weaker
fallback. Model credentials are configured on the worker and never travel in
this document; the guest receives a placeholder://<binding> instead. See
GUARD_POLICY.md.
GET /— a tenant-scoped bare array, filtered by optionalv1/ runs state, ordered by(requested_at DESC,.id DESC) limitdefaults to 50 and is clamped to 1–200. For the next page, pass the last returned run'srequested_atandidasafter_requested_at+after_id; both must be given together or the route returns HTTP 400. The cursor is exclusive. Keep the same state filter and continue until a short or empty page; an exactly full final page requires one additional empty request because this compatible array shape has nonextfield. Consumers expose the same one-page cursor: RustAIecClient::; Pythonlist_runs(state, limit, Option<MatrixCursor>) client.runs.list(state,; CLIlimit, after_requested_at=..., after_id=...) aiec run list --after-requested-at TIMESTAMP --after-id RUN_ID. Python's cursor fields are keyword-only; existing positional state/limit calls remain valid. Python and CLI refuse half cursors before sending.GET /— tenant-scoped detail.v1/ runs/ {id} GET /,v1/ runs/ {id}/ events GET /— the run's history and its attempts.v1/ runs/ {id}/ attempts POST /— stop a run and reclaim its machine. Cancelling a finished run returns it unchanged.v1/ runs/ {id}/ cancel POST /,v1/ eval/ batch POST /,v1/ eval/ repetitions POST /— bounded parallel evaluation.v1/ eval/ matrix batchtakes{requests,and returnsoptions} Vec<Run>;repetitionstakes{request,and returnsrepetitions, options} Vec<Run>;matrixtakes{cells:and returns[{axis, request}], options} {matrix_id,. Every submitted cell appears inrequested_at, max_parallel, results: [{axis, run?, error?}]} results: a cell that was refused before it ran comes back with itsaxisand anerrorand norun, beside the cells that did run, so a partial failure never hides the runs that were executed and billed — and never costs the caller thematrix_idthat is the only way to find them. A matrix must also be keyed throughout or not at all; a partly-keyed matrix is400before anything is admitted, because keyed cells keep their keys across a retry and unkeyed cells do not, and the clash between those two is only visible after the batch has been spent.repetitionsmust be at least 1 andmax_parallelis bounded; every evaluation route needs the sandbox write scope. The Python SDK exposes these asaf.evals.batch(...),.repetitions(...)and.matrix(...), and expands a suite document into thematrixrequest withaf.evals.run_suite(...)and.compare(...);af.evals.comparereports both revisions' outcomes and picks no winner.GET /reads one bounded page back, for a caller that lost the submission response:v1/ eval/ matrix/ {matrix_id} limit(default 50, store-clamped) and the optionalafter_requested_at+after_idcursor, which must be given together. It returns{matrix_id,.cells: [{run, index, axis}], successes, by_axis, next} successesandby_axisdescribe that page only, andnextis the cursor for the following one;cellskeep the labels the cell was submitted under, which are stored on the run at admission.GETneeds the sandbox read scope and is tenant-scoped, so another tenant's matrix is a404, not an empty page. Exposed asaf.evals.matrix_page(...)andclient.eval_matrix_page(...). A run's stored evidence is bounded:setupandvalidationsaccept at most 32 commands each, and the previews kept in the run row are clipped (task 128 KiB, setup and validations 32 KiB each, git evidence 64 KiB), keeping the head and the tail.results.task.output_preview_truncatedandresults.git_evidence_truncatedsay so explicitly and are distinct fromtruncatedand fromok: a command that ran to completion and exited zero is a pass whether or not its output fitted. Full output is collected as a Run artifact when the workload asks for it. Captured stdout, stderr and git evidence are scrubbed of any resolvedworkload.secretsvalues before being stored.results.phase_msdecomposes a run.placementis the whole of getting a machine, andplacement.scheduler,placement.allocation,placement.bootandplacement.workspaceare its parts: reservation, runtimecreate, runtimestart(for Firecracker, rootfs materialization plus VM boot), and workspace preparation (clone or cache import) respectively. A dotted name is a subset of its top-level phase, so a tool that adds up unaccounted client wait must exclude them rather than count the same interval twice.attemptsis the only entry in that map that is a count rather than a duration.
GET / (scope SnapshotsRead) is paged the same way
and returns {snapshots,. Snapshots are retained until one is deleted
and nothing prunes them automatically, so a sandbox that is snapshotted
repeatedly grows that list for as long as it lives. It was the third listing
with the same missing index tie-breaker, after the sandbox list and the two
Guard histories, which is what suggests the shape is now shared deliberately
rather than reproduced.
Run queue#
Every POST / (and matrix/evaluation submission) is admitted to the durable
run queue first and the handler still answers with the settled run, so the response
is unchanged for existing clients. Admission and idempotency are resolved in one
transaction; a full queue is 429 quota exceeded, and a reused
idempotency_key returns the original run without a second execution. The HTTP
connection is not part of run ownership: dropping it stops the waiting, not the
work, and execution continues against the durable run. Limits default to 4
concurrent runs cluster-wide (AIEC_RUN_QUEUE_MAX_ACTIVE), 1024 pending runs
cluster-wide (AIEC_RUN_QUEUE_GLOBAL_PENDING), 128 per tenant
(AIEC_RUN_QUEUE_TENANT_PENDING), a 300 s queue deadline
(AIEC_RUN_QUEUE_TIMEOUT_SECONDS) and a 30 s renewable executor lease
(AIEC_RUN_QUEUE_LEASE_SECONDS). A run that cannot start before its queue
deadline settles failed with Run queue deadline exceeded, and any run whose
executor loses ownership is failed and reclaimed by recovery, never restarted.
A handler waits for the run to settle and for its queue row to reach
finished, so the response carries the persisted cleanup evidence. That second
wait is bounded at 30 s: a teardown that keeps failing repopulates
results.cleanup_failed, which is exactly what keeps the queue row out of
finished, so an unbounded wait would never return. When the bound expires the
run is returned as it stands — outcome known, results.cleanup_failed saying
that teardown is still outstanding — rather than holding the connection open
with no run, no id and no error.
Run secrets#
workload.secrets is a list of secret names, not values. Names are stored
with the run; values are resolved from the operator-configured tenant secret
store (AIEC_RUN_SECRETS_DIR) immediately before each command is executed and
are injected through the command's environment. They are never returned,
persisted, logged or written to a snapshot — see the deployment format in
SECURITY.md.
- A run with no
secretsneeds no store and works unchanged. - A requested name the tenant has no value for fails the run before a machine
is placed, naming the reference (
404/not_found). No placeholder and no empty value is ever substituted. - A name that is not a valid environment variable (uppercase ASCII, digits and underscores, up to 64 characters) is rejected the same way.
- A name present in both
workload.environmentandworkload.secretsis rejected, so a literal in the run document cannot silently shadow a secret. - Names appear in
workload.secretsonGET /. There is no endpoint that returns a run secret's value.v1/ runs/ {id} - Captured command output is scrubbed of resolved values before it is stored, bounded to 8 KiB per stream; error text is scrubbed while keeping its error type.
Run artifacts#
workload.artifacts names paths inside the sandbox — / — collected once the task is done. They are collected into object storage under a server-generated key derived from a digest of the name, so a caller's path never becomes a key and a re-collection overwrites rather than duplicates.
Collection is streamed in binary 64 KiB chunks from the runtime (sandbox → worker → control plane → object store), and the object store writes them incrementally, so no stage holds the whole file or a base64 copy of it in memory. Every chunk read is authorized against the sandbox's active lease generation, and each read revalidates the file's identity (device/inode/size/mtime), so an artifact is never a concatenation of two file generations. A file over the 16 MiB per-file limit is refused and fails the run.
GET /— every collected artifact, each with its recordedv1/ runs/ {id}/ artifacts name(the path as it appeared in the sandbox),object_key,size_bytes,checksum_sha256,content_typeand adownload_url.GET /— the bytes, streamed. The object is read under the checksum recorded at collection and the whole body is verified before the response starts, so a corrupted or overwritten object is refused rather than served; a caller never receives bytes whose digest does not match what was recorded.v1/ runs/ {id}/ artifacts/ {name}
The stored object is the file's own bytes: size_bytes and checksum_sha256 describe the file. The public sandbox artifact API (/) still answers in the same JSON shape with content_base64, but it too encodes from the bounded verified stream. download_url percent-encodes the name, so an absolute or nested name round-trips exactly; a hand-written path still has to name an artifact this run collected, and cross-tenant reads are 404.
Collection failures fail the run rather than returning a shorter list than was asked for. A requested artifact that cannot be read, exceeds the limit, changes while it is being read, or cannot be stored sets failure_reason naming the artifact, and the run settles failed. Artifacts collected before the failure are still recorded and downloadable, because a failed run is the one whose evidence somebody opens.
Operational endpoints: /, /, /.
API keys#
Three routes, and only three: GET / lists the caller's own tenant's
keys; POST / mints one; DELETE / revokes one. There is no
rotate route and there is no POST / — revocation is the
DELETE verb on the key itself.
Key management is the one place where omitting a field grants authority, so
it is worth being explicit. POST / with no scopes is a request for
the default set — sandboxes:, sandboxes:, snapshots:,
snapshots: — and that is a grant, not a shortcut. Every scope that would
be granted, defaulted or explicit, must be held by the calling credential, so a
key carrying only sandboxes: is refused both an omitted scopes field and
an explicit ["sandboxes:. The rule is the same on both paths on purpose:
the default is a way of asking for four scopes, not a way around needing them.
Revocation follows the same principle, because destroying a credential is a use
of authority. To revoke a key, the caller must hold every scope the target
holds. A sandboxes: key cannot revoke a key that can write, and neither
can it revoke a tenant admin key. A key from another tenant, or one that does
not exist, is 404 — a key's existence is not disclosed across a tenant
boundary. Revoking a key revokes it immediately; there is no grace period.
Status codes#
400 invalid JSON/path/argv/image; 401 missing/invalid/expired/revoked key; 403 missing scope or cross-tenant access (cross-tenant resources are not disclosed); 404 tenant-scoped missing resource; 409 invalid state/race; 413 upload/resource limit too large; 422 semantic validation; 429 PostgreSQL tenant quota exceeded; 500 internal; 503 runtime/dependency unavailable.