AIec

Design

5 min readdocs/DESIGN.md

AIec is an implementation of the ideas in DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale (arXiv:2609.22978), reduced to what a small, honest implementation can actually stand behind.

This document records which mechanisms came from the paper, which are independent engineering, and where AIec deliberately does something different.

The problem the paper names#

Agentic training and evaluation need sandboxes that:

DSec's conclusion is that this needs an elastic execution platform, not a single sandbox runtime. Two consequences shape AIec:

  1. One contract, many backends. A caller should not have to care whether a sandbox is a microVM, a container, or a hosted provider's machine.
  2. Ownership must be fenced. With sandboxes moving between workers, a worker that has lost its lease must be unable to touch a sandbox it no longer owns.

Mechanism mapping#

DSec mechanism AIec implementation
Unified SDK across FnCall / container / microVM / full-VM backends SandboxRuntime in crates/aiec-core/src/runtime.rs, implemented by Firecracker, the hosted E2B adapter, and Docker
Heterogeneous, capability-aware placement RuntimeCapabilities advertised per worker, persisted, and matched by the scheduler before placement
Bursty arrival and resource admission Per-tenant quotas enforced inside the placement transaction, per-tenant rate limiting, and a global execution budget
Sandbox lifecycle with recoverable state Leased ownership with a monotonic fencing generation; workspace archives in shared object storage, restored on a new owner
Image corpora with limited reuse Content-addressed image references, signed manifests, digest verification before a guest boots

Where AIec differs#

These are choices, not omissions of the paper.

Guard: what the paper does not name#

The paper's isolation story ends at the guest boundary. That is the right scope for training a model, where the model is the thing you are producing. For deploying an agent that acts on a developer's machine, the interesting failure is one step further in: a sandbox whose guest has been persuaded to want something the operator never agreed to.

Guard answers that by moving the decisions out of the machine. A sandbox is still a microVM with its own kernel; what changes is that the answer to "what may this reach, with whose credentials, and recorded where" is computed on the worker, from a policy the guest cannot read, using secrets the guest never holds, into a journal the guest cannot rewrite. The consequence worth stating is the failure mode: compromising the guest does not grant new capability, because capability was never expressed in terms the guest could influence.

Three design choices follow from that, and each one is a refusal rather than a capability:

The full contract, the operator settings and the phase-by-phase status are in GUARD_POLICY.md, docs/guard-plan.md and docs/openshell-compatibility.md.

Fencing, in detail#

A sandbox is owned through a lease carrying a generation that only ever increases. Every state-changing operation is checked against the current owner and generation before the runtime is touched:

worker A owns generation N
  → A dies
  → its lease expires
  → the control plane reassigns to worker B at generation N+1
  → B reconstructs the workspace from the newest archive
  → if A returns, its generation-N operations are rejected

This is why the lease lives in the database rather than in a worker's memory: a restarted worker that still believes it owns a sandbox must be told otherwise.

What AIec does not do#

Reference#