Ecosystem
Most confusion about tools in this space comes from comparing things that occupy different layers. E2B and SWE-bench are not competitors. MCP and a task standard are not competitors. Putting them in one ranking is a category error, and this page exists to stop it.
The layers
| Layer | Question it answers | Examples |
|---|---|---|
| Execution substrate | Where does this code actually run? | E2B, Daytona, Modal, AIec, Firecracker, Docker |
| Local process confinement | How do I sandbox a process on my own machine? | Bubblewrap, Anthropic Sandbox Runtime |
| Evaluation runner | How do I score an agent on a set of tasks? | Inspect AI |
| Task standard | How is one task specified so any harness can run it? | METR Task Standard |
| Benchmark | Which tasks are the test set, and what is the score? | SWE-bench |
| Protocol | How do tools and models talk to each other? | MCP |
AIec occupies the execution-substrate and orchestration layers, and provides an evaluation surface on top of its own substrate. It is not a benchmark, not a task standard, and not a protocol.
Project by project
E2B
Layer: execution substrate. E2B provides cloud-hosted secure sandboxes for AI agents, with SDKs for creating and running code in an isolated environment.
The boundary: both AIec and E2B answer "where does this code run, and what can it reach". The difference is where the machines live. E2B is a hosted service; AIec runs on hardware you operate, which means the isolation boundary is one you can inspect, and there is no per-run billed dependency on someone else's capacity. AIec is not a hosted offering and does not compete for that use case.
Daytona
Layer: execution substrate. Daytona provides development environments and infrastructure for running code.
The boundary: a development-environment lifecycle overlaps AIec's sandbox lifecycle, but AIec's concern is short-lived machines for untrusted agent-written code rather than long-lived workspaces a human works in. A durable developer workspace and an ephemeral agent sandbox want opposite retention policies.
Modal
Layer: execution substrate, with its own serverless scheduling model. Modal runs functions and containers on managed GPU and CPU capacity.
The boundary: Modal is capacity you rent per unit of work. AIec's capacity is machines you already own, and it refuses a placement rather than spilling work elsewhere. If the question is "who provides compute when my cluster is full", the answers differ by design, not by quality.
Inspect AI
Layer: evaluation runner. Inspect AI provides a framework for building evaluations, with solvers, scorers and a logging and inspection interface.
The boundary: Inspect AI scores agents; AIec puts them somewhere to run. They compose — an Inspect evaluation can target sandboxes as its execution environment — and AIec includes its own evaluation surface so that a run's environment, policy and artifacts are recorded with its score. The two are not alternatives.
METR Task Standard
Layer: task standard. It specifies how a single task should be expressed so that different harnesses can run and compare it.
The boundary: a task standard says how to describe work; AIec says where that work runs and under what policy. AIec is not a task standard and does not define a competing one.
SWE-bench
Layer: benchmark. A fixed set of real software issues, each with a test that determines whether a patch fixed it, and a leaderboard of results.
The boundary: SWE-bench is the test set; AIec is the environment the agent works in while attempting it. Running an agent against SWE-bench needs somewhere to run it — that is the layer AIec provides. AIec is not a benchmark and publishes no leaderboard position.
MCP
Layer: protocol. The Model Context Protocol is how clients and tools expose capabilities to models.
The boundary: AIec ships an MCP server so an agent can drive sandboxes through a standard interface rather than a bespoke one. MCP is the wire format; it says nothing about where the sandbox lives or how it is isolated. AIec implements a server for the protocol, it does not define or own the protocol.
Anthropic Sandbox Runtime
Layer: local process confinement. A tool wrapping Bubblewrap and network filtering for running commands locally under constraints.
The boundary: this is the closest neighbour in spirit, and the difference in scope is the point. Local confinement runs a process on your own machine with a namespace and seccomp profile. AIec distributes machines across workers, owns their lifecycle and leasing, enforces policy from outside the guest, and records evidence that survives the machine. AIec uses Bubblewrap as one of its own runtimes rather than competing with a wrapper around it.
AIec
Layer: execution substrate and orchestration, with evaluation on top. It places, leases, isolates and destroys machines; governs what a run can reach from outside the guest; and retains the artifacts and policy evidence after the machine is gone.
Where AIec is not the right tool
- If you want a hosted service. There is no managed AIec and no sign-up. Capacity comes from your hardware.
- If you want to run one command locally with tight sandboxing. Use Bubblewrap directly. AIec's machinery earns its cost across many machines with lifecycle, tenancy and evidence requirements.
- If you want a benchmark or a leaderboard. Use one that exists. AIec is not a test set.
- If you want a protocol. Use MCP. AIec implements a server for it.
Most of the projects above are useful and, where they overlap, often complementary. A substrate can host an evaluation runner; an evaluation runner can use a task standard; a task standard can describe benchmark tasks. The genuinely contested ground here is narrow — where a substrate owns its own policy enforcement and evidence retention — and that is the only place worth arguing about.