AIec

Local MCP

Point any MCP-capable agent at your own AIec cluster and it gets a disposable computer per task — a real machine, created and destroyed on demand.

Everything below runs on your hardware. The server is local-only by construction: it refuses to start against a remote AIec endpoint, and it cannot reach a hosted provider, so a workload cannot escape the cluster you control.

Connecting a client

URL

http://127.0.0.1:8765/mcp

Streamable HTTP. The server binds to loopback and will refuse any other address unless you opt in.

Authorization

Bearer <token>

Generated on first run and stored at ~/.config/aiec/mcp-token, mode 0600.

export AIEC_LOCAL_API_URL=https://127.0.0.1:18443
export AIEC_LOCAL_API_KEY=af_live_...   # a key you created on your own control plane
./target/release/aiec-mcp

What the agent gets

Lifecycle

aiec_create_sandbox, aiec_list_sandboxes, aiec_get_sandbox, aiec_destroy_sandbox

Destroy waits for cleanup to be confirmed and is safe to call twice.

Working in a machine

aiec_exec, aiec_read_file, aiec_write_file, aiec_list_files

Commands run inside the sandbox with bounded output and a timeout. Nothing runs on the host that serves MCP.

Repository work

aiec_prepare_repo, aiec_run_repo_task

Clone, run a task, validate, collect the diff, then destroy the machine — on the failure path too.

Evaluating an agent

aiec_test_omp, aiec_compare_omp

Run a coding agent in a clean sandbox, or compare two revisions against the same target, task and validations.

A loop an agent can run unattended:

{"name": "aiec_create_sandbox", "arguments": {"image": "aiec-coding:latest"}}
{"name": "aiec_exec", "arguments": {"sandbox_id": "<id>", "command": ["git", "status"]}}
{"name": "aiec_destroy_sandbox", "arguments": {"sandbox_id": "<id>"}}

Comparing two revisions of an agent

"I changed the compaction implementation. Does it actually improve things?"

{"name": "aiec_compare_omp", "arguments": {
  "omp_repo": "https://github.com/can1357/oh-my-pi",
  "baseline_ref": "main",
  "candidate_ref": "feature/compaction-v2",
  "target_repo": "https://github.com/me/fixture",
  "task": "add a regression test for truncated output",
  "repetitions": 3
}}

Each side gets its own clean machine, same target, same task, same validations, same resources. You get successful and failed runs, validation passes, wall time, exit codes, changed files and diff sizes — measurements, not a verdict. All sandboxes are cleaned up afterwards.

What it will not do

Full reference: docs/MCP.md in the repository.