AIec

REST API

13 min readdocs/API.md

Base path /v1. Except health/metrics, send Authorization: Bearer af_live_.... JSON uses snake_case; errors are {"error":{"code":"...","message":"...","request_id":"..."}}. Every response includes x-operation-id; clients may supply a UUID in that header, otherwise the server generates one. This correlation header is not yet propagated into structured logs.

Sandboxes#

The Python SDK exposes the same size-guarded partial path as Sandbox.upload_artifact, download_artifact, and delete_artifact; listing remains unsupported, and the production S3 path remains unverified.

Runs#

Both Guard history listings are paged the same way, for the same reason. A proposal row is never reclaimed, and a decided tool approval is deliberately retained, so a sandbox left running grows both without bound. GET /v1/sandboxes/{id}/guard/proposals (scope GuardRead) and GET /v1/sandboxes/{id}/guard/tool-approvals (scope GuardApprove) take the same limit (default 50, store-clamped to 200) and the same paired after_created_at + after_id cursor, and return {proposals, next} and {approvals, next} respectively, ordered created_at DESC, id DESC, with next non-null exactly when another page follows. The id half of the cursor is not decoration: created_at is not unique, and a keyset that cannot break a tie either re-reads or skips the rows sharing a timestamp.

resources.guard selects an out-of-guest policy: either {"policy_template": "no-network" | "model-only" | "model-plus-allowlist" | "read-only-api"} with optional model_endpoint and allowlist inputs, or a full {"policy": {...}} document. It replaces resources.network rather than accompanying it — two descriptions of one egress would be two paths to the internet — and it is enforced by the worker's Guard gateway, not by the guest. A guarded run is placed on a microVM runtime; there is no weaker fallback. Model credentials are configured on the worker and never travel in this document; the guest receives a placeholder://<binding> instead. See GUARD_POLICY.md.

GET /v1/sandboxes/{id}/snapshots (scope SnapshotsRead) is paged the same way and returns {snapshots, next}. Snapshots are retained until one is deleted and nothing prunes them automatically, so a sandbox that is snapshotted repeatedly grows that list for as long as it lives. It was the third listing with the same missing index tie-breaker, after the sandbox list and the two Guard histories, which is what suggests the shape is now shared deliberately rather than reproduced.

Run queue#

Every POST /v1/runs (and matrix/evaluation submission) is admitted to the durable run queue first and the handler still answers with the settled run, so the response is unchanged for existing clients. Admission and idempotency are resolved in one transaction; a full queue is 429 quota exceeded, and a reused idempotency_key returns the original run without a second execution. The HTTP connection is not part of run ownership: dropping it stops the waiting, not the work, and execution continues against the durable run. Limits default to 4 concurrent runs cluster-wide (AIEC_RUN_QUEUE_MAX_ACTIVE), 1024 pending runs cluster-wide (AIEC_RUN_QUEUE_GLOBAL_PENDING), 128 per tenant (AIEC_RUN_QUEUE_TENANT_PENDING), a 300 s queue deadline (AIEC_RUN_QUEUE_TIMEOUT_SECONDS) and a 30 s renewable executor lease (AIEC_RUN_QUEUE_LEASE_SECONDS). A run that cannot start before its queue deadline settles failed with Run queue deadline exceeded, and any run whose executor loses ownership is failed and reclaimed by recovery, never restarted.

A handler waits for the run to settle and for its queue row to reach finished, so the response carries the persisted cleanup evidence. That second wait is bounded at 30 s: a teardown that keeps failing repopulates results.cleanup_failed, which is exactly what keeps the queue row out of finished, so an unbounded wait would never return. When the bound expires the run is returned as it stands — outcome known, results.cleanup_failed saying that teardown is still outstanding — rather than holding the connection open with no run, no id and no error.

Run secrets#

workload.secrets is a list of secret names, not values. Names are stored with the run; values are resolved from the operator-configured tenant secret store (AIEC_RUN_SECRETS_DIR) immediately before each command is executed and are injected through the command's environment. They are never returned, persisted, logged or written to a snapshot — see the deployment format in SECURITY.md.

Run artifacts#

workload.artifacts names paths inside the sandbox — /workspace/report.txt — collected once the task is done. They are collected into object storage under a server-generated key derived from a digest of the name, so a caller's path never becomes a key and a re-collection overwrites rather than duplicates.

Collection is streamed in binary 64 KiB chunks from the runtime (sandbox → worker → control plane → object store), and the object store writes them incrementally, so no stage holds the whole file or a base64 copy of it in memory. Every chunk read is authorized against the sandbox's active lease generation, and each read revalidates the file's identity (device/inode/size/mtime), so an artifact is never a concatenation of two file generations. A file over the 16 MiB per-file limit is refused and fails the run.

The stored object is the file's own bytes: size_bytes and checksum_sha256 describe the file. The public sandbox artifact API (/v1/sandboxes/{id}/artifacts/{name}) still answers in the same JSON shape with content_base64, but it too encodes from the bounded verified stream. download_url percent-encodes the name, so an absolute or nested name round-trips exactly; a hand-written path still has to name an artifact this run collected, and cross-tenant reads are 404.

Collection failures fail the run rather than returning a shorter list than was asked for. A requested artifact that cannot be read, exceeds the limit, changes while it is being read, or cannot be stored sets failure_reason naming the artifact, and the run settles failed. Artifacts collected before the failure are still recorded and downloadable, because a failed run is the one whose evidence somebody opens.

Operational endpoints: /health, /ready, /metrics.

API keys#

Three routes, and only three: GET /v1/keys lists the caller's own tenant's keys; POST /v1/keys mints one; DELETE /v1/keys/{id} revokes one. There is no rotate route and there is no POST /v1/keys/{id}/revoke — revocation is the DELETE verb on the key itself.

Key management is the one place where omitting a field grants authority, so it is worth being explicit. POST /v1/keys with no scopes is a request for the default set — sandboxes:read, sandboxes:write, snapshots:read, snapshots:write — and that is a grant, not a shortcut. Every scope that would be granted, defaulted or explicit, must be held by the calling credential, so a key carrying only sandboxes:read is refused both an omitted scopes field and an explicit ["sandboxes:write"]. The rule is the same on both paths on purpose: the default is a way of asking for four scopes, not a way around needing them.

Revocation follows the same principle, because destroying a credential is a use of authority. To revoke a key, the caller must hold every scope the target holds. A sandboxes:read key cannot revoke a key that can write, and neither can it revoke a tenant admin key. A key from another tenant, or one that does not exist, is 404 — a key's existence is not disclosed across a tenant boundary. Revoking a key revokes it immediately; there is no grace period.

Status codes#

400 invalid JSON/path/argv/image; 401 missing/invalid/expired/revoked key; 403 missing scope or cross-tenant access (cross-tenant resources are not disclosed); 404 tenant-scoped missing resource; 409 invalid state/race; 413 upload/resource limit too large; 422 semantic validation; 429 PostgreSQL tenant quota exceeded; 500 internal; 503 runtime/dependency unavailable.