Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

26. Why it is built this way

Some choices in Zygo look strange until you know the reason: why the vm backend has no warm functions, why there is no Deno agent, why an upgrade throws the warm sandboxes away. Each of these was written down as an Architecture Decision Record (ADR): a short, dated note of a question, the answer, and what would change the answer. This chapter explains the nine ADRs in plain words; each section links to the full record in adr/.

The nine decisions at a glance

ADRThe questionThe answer, in one line
0001Who is Zygo for?Programs that embed it to run many scripts — not a person typing commands.
0002Should vm and gvisor have warm functions?No. Warm functions are an ns feature.
0003Should there be Deno and Bun agents?No, until an embedder asks with a measurement.
0004Can an upgrade keep the warm sandboxes?No. An upgrade drains, restarts and re-warms.
0005One warm zygote per script version? What evicts it?Yes for heavy scripts, a runtime pool for light ones; timeouts evict.
0006What does mem bound in a warm function?One request, on its own cgroup — never the function as a whole.
0007Where does an Elixir client live, and what does it need?Here, as sdk/elixir; Mint and NimblePool; one error with a kind.
0008May a sandbox listen on its own loopback?A function may; a runtime pool, shared between tenants, may not.
0009May what the API serves reach a private address?Only when zygo api --allow-private-net says so; never a request body.
  how the nine decisions depend on each other
  ─────────────────────────────────────────────────────────────────
   0001  the product is the warm path, for an embedder
     │
     ├──► 0002  keep one warm path (ns), do not build three
     ├──► 0003  keep the agent list small enough to test fully
     ├──► 0004  restarts re-warm; keep one copy of the state
     ├──► 0005  how a real embedder should use warm paths
     ├──► 0006  one tenant's request cannot take the others down
     ├──► 0007  an embedder reaches it through an SDK in its own language
     ├──► 0008  a function may listen; a pool shared by tenants may not
     └──► 0009  the operator, never a request, opens private addresses
  ─────────────────────────────────────────────────────────────────

Each ADR ends with the facts that would reopen it. None of them is “forever”. They say “not until this is true”.

How to read an ADR

An ADR is a historical record. It is written once, at the time of the decision, and not rewritten later. If a decision changes, a new ADR replaces the old one. So an ADR can mention files or plans that have since moved. The sections below are today’s plain-English summary; the ADR itself is the exact wording.

ADR 0001: Zygo is the embedded script runtime

The question

Zygo can be described in two ways. One is “a faster Docker”: one static binary, no daemon, no root, OCI images, a sandbox in about 12 ms instead of 300–1000 ms. That is true, but the field is crowded — kern, nono, microsandbox and Docker’s own sandboxes all compete there. The other is “the runtime a workflow engine embeds”: a program with ten thousand scripts in a database, each run a few times a minute. The question was which of the two is the product.

The decision

The product is the second one. The target user is an embedder: a workflow engine (Windmill, n8n, Temporal, Kestra), a SaaS running plugins its customers wrote, or an agent platform running generated code. Four things follow from that, in order of how often they come up:

  1. The warm path is the product. A change that costs the warm path milliseconds needs a much better reason than one that costs the one-shot path the same.
  2. The API is the surface, not the CLI. The CLI stays for debugging and demos.
  3. Tenants are strangers to each other, so isolation between tenants gets the effort.
  4. One process per worker, one host. No scheduler, no control plane, no Helm chart — the embedder has those already.

Why

Nobody else serves the “ten thousand scripts, each run often” shape. A warm fork is the tool built for it: each request is a fork() of a process that never served a request, so it is as clean as a fresh container and as cheap as a fork. That is the one part another project would have to rebuild rather than just make faster. The one-shot sandbox still matters, but mostly because the warm sandbox is built out of it.

What it costs

Some things are put aside on purpose, so nobody reopens them by accident: services, ports, compose files and restart policies; warm functions on vm and gvisor (see ADR 0002); the macOS shim’s latency; Windows; and a hosted service, which would compete with the very users Zygo is for. Work that is good for a developer at a terminal but bad for an embedder now loses. The macOS hop is the standing example: a developer notices it, a production embedder never sees it.

The gates, and how they were met

The ADR sets a gate for each phase of the roadmap, so that a phase cannot be declared done by whoever did the work.

  phase                        done when …                                    status
  ─────────────────────────────────────────────────────────────────────────────────
  0 prove the wedge            warm fork ≥ 10× under the best one-shot runner  met
  1 runtime zygotes            1 000 scripts, < 5 ms, memory flat              met
  2 embedder API               a plugin host needs only the HTTP API
  3 runtimes                   Python and JavaScript through one API
  4 deployability              kubectl apply → green probe + agent test passes
  5 hardening                  no "untested" rows for multi-tenant claims
  6 integrations               one outside project runs it in production
  ─────────────────────────────────────────────────────────────────────────────────

Phase 0 was a real gate: without the 10× ratio, the ADR would be wrong, not early. It passed at 25× on the Lima VM — and at 60× on Docker Desktop’s VM before the bytecode layer, and 100× on a Raspberry Pi in an older run that chapter 25 has not repeated. Phase 1 passed on the kernel it was set on: on Docker Desktop’s Linux 5.10 VM even the slowest 1 in 100 calls took 3.2 ms, and a thousand scripts in one zygote used 29.4 MB. On the Lima VM’s Linux 6.8 the same 1 in 100 is 11.4 ms as the kernel comes, and 3.3 ms once zygo doctor --fix has turned on favordynmods; a warm function with no pool at all shows the same tail, so it is the kernel’s cgroup move, not the pool. Chapter 25 explains the tail and has both measurements.

What would reopen it

The ADR names no single trigger; it stands as long as its Phase 0 gate holds. If the warm fork stopped being many times cheaper than the best one-shot runner on an import-heavy script, the idea behind the product would be wrong. Full ADR 0001.

ADR 0002: Warm functions stay on ns

The question

Zygo has three isolation backends behind one flag: ns (namespaces, cgroups, seccomp and Landlock), gvisor (a kernel written in user space) and vm (libkrun, a real hardware boundary). All three run one-shot sandboxes from the same spec. Only ns runs warm functions. The question was whether the other two should get warm functions too, or whether the gap was just unfinished work.

The decision

Warm functions are an ns feature. vm and gvisor run one-shot sandboxes and refuse warm modes with a reason. That refusal is the design, not a gap. Networking on vm is refused for the same reason. The error messages point at this ADR instead of saying “not yet”.

Why

Both gaps come from how the backends are built, not from missing time:

  • The agent gets its control socket as an inherited file descriptor. An OCI runtime such as runsc closes everything except stdin, stdout and stderr, and a guest VM inherits nothing from the host at all. gvisor would need runsc exec instead of setns; vm would need the protocol carried over vsock and a supervisor inside the guest.
  • Networking on vm would need the VM monitor inside Zygo’s own network namespace, and the allowlist applied to a guest interface. That is a second copy of the network code with none of the first copy’s tests.

Each is weeks of work. Worse, each makes a second warm path that must be kept correct, measured and defended — next to the one the whole product rests on.

                      one-shot     warm function    network
  ─────────────────────────────────────────────────────────────
  ns                  yes          yes              yes
  gvisor              yes          refused          —
  vm                  yes          refused          refused
  ─────────────────────────────────────────────────────────────

What it costs

--isolation vm is a hardware boundary for work that fits a one-shot sandbox: an untrusted build, a single tool call, a job with an input and an output. It is not for a warm function serving many requests. An embedder who needs a hardware boundary per tenant is not served by Zygo today; chapter 10 points at an alternative. The effort goes instead into hardening ns, which stays one kernel away from the host (chapter 23).

What would reopen it

An embedder who asks for a hardware boundary per tenant, with a workload that can afford about 100 ms of boot per request. That is a different product shape from the warm fork. It should be thought through as its own thing, not bolted onto this backend just because the flag already exists. Full ADR 0002.

ADR 0003: No Deno or Bun agent until an embedder asks

The question

Zygo ships two runtime agents, Python and Node. A third-party agent can be added with agent = { agent = "/path/in/sandbox" }. Deno and Bun are the obvious next two: both are popular, both start fast, and both would look good in a table. The question was whether to write agents for them.

The decision

No Deno or Bun agent. Anyone who wants either has two supported paths. One is a warm-exec pool, which works today with no agent at all, because both start in a few milliseconds:

[runtime.deno]
image = "denoland/deno:alpine"
cmd   = ["deno", "run", "--allow-none"]

The other is to write their own agent against the protocol and check it with zygo agent test (chapter 18).

Why

An agent is cheap to write and expensive to keep. Each one is a fork boundary with four rules that must be exactly right: the child never returns to the parent’s loop, nothing runs before GO, a broken frame is reported and not fatal, and every EXEC gets exactly one answer. Each rule has been broken at least once in this repository, by someone who knew the language well. Each agent also needs its own answer to the per-request seccomp filter; Node’s fallback was once stricter than the filter it replaced, which the seccomp matrix caught. And each agent must sit in make conformance, the seccomp matrix and the image matrix — a test nobody runs is only a claim.

Meanwhile Deno and Bun both run JavaScript, which the Node agent already serves. Someone who asks for Deno usually wants its permission model, its standard library or deno.json, not “something other than V8”.

What it costs

The comparison table says two agents, not four — the honest number. The test matrices stay small enough to run on every change. A Deno user starts with a three-line cmd and no protocol. Each request is slower than a fork from a warm heap by the cost of an execve, not by the cost of an interpreter start-up. The warm-exec pool gives up streaming, progress(), workspaces and per-request tenant limits.

What would reopen it

Any one of three facts — none of them a guess about the future:

  • An embedder asks, with a workload where the warm-exec pool’s per-request execve is measurably too slow. The measurement is the argument, not the runtime’s popularity.
  • A dependency set needs it: supporting deno.json or bun.lockb in POST /deps, which is smaller work than an agent.
  • Someone writes an agent and it passes zygo agent test, including the child-filter checks. The project would rather link to it than rewrite it.

Full ADR 0003.

ADR 0004: A supervisor upgrade re-warms; there is no --reexec

The question

Upgrading Zygo replaces the supervisor process, and its warm sandboxes go with it. Could the new binary exec over the old one and keep them, so an upgrade costs nothing? exec is the right tool to ask about: it keeps the process id, so the sandboxes’ init processes stay children of the same process, and nothing dies just because the binary changed.

The decision

No zygo api --reexec. A supervisor upgrade is a restart: drain, exit, start, re-warm. min_warm and a rolling update with maxUnavailable: 0 make it invisible to callers, and both already exist (chapter 16).

  an upgrade, with two replicas and maxUnavailable: 0
  ──────────────────────────────────────────────────────────────────────
  old supervisor   serving ████████████ POST /drain ▓▓▓ finish ▓▓ exit
  new supervisor              start ░░ re-warm ░░ ready ████████ serving
                                        ~150–185 ms per Python zygote
  callers see:     no dropped request; at worst, slower ones while warm-up runs
  ──────────────────────────────────────────────────────────────────────

Why

What would have to cross the exec is much more than process ids:

  • Every file descriptor, on purpose. Each warm function holds an agent socket; each sandbox holds seven namespace descriptors, a /run/secrets descriptor and a cgroup descriptor. All are close-on-exec (CLOEXEC) by design — once after a real bug, where thirteen descriptors leaked into every tenant program. A hand-over turns “everything closes unless something says otherwise” into “everything closes unless this list says otherwise”, and the list changes per function, per sandbox and per release.
  • All the state that is not a descriptor: every function’s spec, counters, logs, secrets, script leases, queue counts and idle clocks. That means a second, versioned copy of the whole supervisor state, which could go wrong silently.
  • The requests in flight. Their answers are owed to client connections tracked by a thread that exec destroys. Keeping them means handing over the client sockets, each request’s cgroup, child and deadline, and rebuilding the router. Anything less drops requests — the one thing the feature was for.

What it costs

A Python zygote re-warms in about 150–185 ms on a Raspberry Pi 5, the supervisor’s start included, and in 34 ms on the Lima VM with a supervisor already running (chapter 25). When this was decided, before the bytecode layer, it was about 500 ms (502 ms for a pool, 508 ms for a function, from tests/linux/api_driver.py). A Node zygote is ready in 15–43 ms. So an upgrade costs min_warm × warm-up of cold pool per replica, once. A single-replica deployment has a short window where requests are slow, not failed; two replicas remove it, and the Kubernetes example uses two. In return, descriptors stay CLOEXEC with no exceptions, and Zygo keeps only one copy of its state: the running one.

What would reopen it

  • A warm-up of tens of seconds, not a fraction of one. Even then, the better fix may be to make that warm-up faster.
  • A deployment that cannot have two replicas, with a latency budget a cold start breaks. --reexec is one answer; a second supervisor on the same host behind a load balancer is another, and needs nothing new.

What does not reopen it: wanting upgrades to be free. They are already free of dropped requests, which is the part that matters. Full ADR 0004.

ADR 0005: One warm zygote per script version, and what evicts it

The question

The shape that forced the question is a multi-tenant application: tenant code that changes whenever someone presses Save, hundreds of projects, each with its own mounts, network allowlist and per-run secrets. Such a product starts with zygo run, a sandbox per event, and asks three things. Is one warm zygote per script version the intended shape? What evicts warm zygotes when a server has four hundred projects? And can secrets and mounts vary under one warm zygote?

What was measured

A hundred distinct Python handlers on the Lima VM (Ubuntu 24.04, kernel 6.8, aarch64, 2 vCPU, 3.8 GiB), python:3.12-slim:

per warm script100 scripts
memory (RSS)21.4 MB2 143 MB
memory, shared pages split fairly (PSS)11.3 MB1 141 MB
time to warm109 ms10.9 s
The same hundred scripts in one runtime pool
memory, whatever the count29.5 MB, one zygote
what one more script adds0 kB
second call over the API: usually / 1 in 1001.95 ms / 2.38 ms
a warmed function on the same host: usually / 1 in 1001.36 ms / 1.57 ms
the pool’s extra cost+0.5 ms usually, +0.8 ms for the slowest 1 in 100

So about 300 warm Python scripts fit in 4 GB on that VM, and a thousand would need 11 GB. Chapter 25 explains RSS and PSS.

The decision, part 1: which shape

One warm zygote per script version is right when the script’s imports are worth paying once — an ML model, a large client library. It pays those once and about 1.4 ms per call, and costs 11 MB while warm. A runtime pool is right when the script is a few lines over the standard library, like most such scripts. It pays about 0.5 ms more per call and nothing to stay resident, because the script is not kept; it arrives with each request. (Chapter 25 puts the pool’s cost at +0.47 ms on the Lima VM and +0.46 ms on Docker Desktop’s; the ADR quotes 0.6 ms, from its own earlier run.)

  does the script import something expensive?
  ─────────────────────────────────────────────────────────────
     yes ──► one warm zygote per script version
             pays the imports once · ~1.4 ms a call · ~11 MB warm
     no  ──► a runtime pool
             +0.5 ms a call · 0 kB per extra script
  ─────────────────────────────────────────────────────────────

The decision, part 2: what evicts

Eviction is idle_timeout and cold_after, per function, and nothing else in the product. A version nobody called for idle_timeout (default ten minutes) is paused: frozen, still in memory, one write to wake. Past cold_after (default an hour) it is cold: the sandbox is dropped, only the spec is kept, and the next call pays the 109 ms plus imports. An LRU list of warm scripts is the consumer’s own policy on top, choosing which versions to serve at all. max_warm is a different knob: a pool’s ceiling on its own zygotes under load. When the host is full, the answer is a 429, not a swap storm.

  a script version's life
  ─────────────────────────────────────────────────────────────────────
  warm ──(no call for idle_timeout, 10 min)──► paused ──(cold_after, 1 h)──► cold
   ▲                                             │                          │
   └──────────── one write to wake ◄─────────────┘                          │
   └──────────── 109 ms + imports to warm again ◄───────────────────────────┘
  ─────────────────────────────────────────────────────────────────────

The decision, part 3: secrets and mounts

Per-run secrets can vary under one zygote. Their names are declared on the function; their values arrive with the request and exist as /run/secrets/<name> only while it runs, never in the zygote. Mounts and the network allowlist cannot. They are the sandbox’s namespaces, built once at warm-up. So one script version used by two projects with different mounts is two functions, named {project}-{digest} — which matches the adopter’s own model, where a script belongs to a project.

What it costs, and what would reopen it

A busy consumer keeps as many warm zygotes as were called in the last idle_timeout, and zygo ps shows how many. The density benchmark must be run again when the interpreter, the image or the kernel changes. Left open for a later ADR: a global warm budget (max_warm_total) that evicts the least recently called function when the host nears its memory limit. Nothing measured says it is needed before idle_timeout and a 429 do their job; a consumer who shows that need would reopen it. Full ADR 0005.

ADR 0006: The memory limit is each request’s, not the function’s

The question

A warm function or a runtime pool runs several requests at once in one sandbox: a zygote, and one process per request under it. Its mem limit was written on the function’s cgroup, the group that holds the zygote and every request together. So mem was one shared budget, and the kernel’s “kill the whole group” setting sat at that level too. In a pool at mem = 256M, one request that asked for 2 GB was killed — and so were the zygote and the requests sleeping beside it, in the same moment. In a runtime pool those other requests can belong to other tenants. The book promised the opposite: that a request has its own group and can be killed alone. The question was where mem should really be written.

The decision

mem goes on the leaves — the smallest groups at the bottom of the tree — and nowhere above them. Each request’s own cgroup gets mem, with the group kill turned on there. The zygote’s leaf gets the same limit, so the warm process is bounded on its own. The function’s cgroup keeps the limits that really are one budget for the whole function: processes, CPU and swap. Its memory limit is set to “none” on purpose, so a folder left over from an older Zygo does not keep the old shared limit.

  BEFORE                                  AFTER
  ┌─ function: mem, group kill ──────┐    ┌─ function: pids, cpu, swap ──────┐
  │ ┌────────┐ ┌─────────┐ ┌───────┐ │    │ ┌────────┐ ┌─────────┐ ┌───────┐ │
  │ │ zygote │ │ req (a) │ │ req b │ │    │ │ zygote │ │ req (a) │ │ req b │ │
  │ └────────┘ └─────────┘ └───────┘ │    │ │  mem   │ │   mem   │ │  mem  │ │
  └──────────────────────────────────┘    │ └────────┘ └─────────┘ └───────┘ │
  (a) goes over: all three die            └──────────────────────────────────┘
                                          (a) goes over: only (a) dies

What it costs

A tenant’s narrower limits still go on the request’s cgroup, in place of the function’s numbers. A whole function may now use up to (concurrency + 1) × mem in each sandbox, not mem, so a host sized as “functions × mem” was sized for the old rule; chapters 13 and 20 say so. Memory a request shares with the zygote from the fork stays counted on the zygote, so only what a request allocates after it starts counts against its own mem. No extra ceiling was added above the leaves: it could only fire on memory the zygote holds, and a kill there would again reach the wrong process. make verify-oom-linux checks the promise, for the Python and the Node agent: a request that goes over dies, and the requests beside it finish.

What would reopen it

  • A kernel that moves a process’s memory charge with it when it changes cgroup. Then a ceiling above the leaves would be exact, and worth adding.
  • An embedder who wants mem to mean the whole function’s budget again, with a measurement of what the per-request shape costs them.

Full ADR 0006.

ADR 0007: A third SDK, in Elixir, in this repository

The question

The first product built on Zygo from Elixir ran each script by starting the zygo binary, and about a thousand lines of it worked around that: stdin through a shell, output order lost on macOS, a warm-function table behind a global lock, errors guessed from their text. The HTTP API already answers all of it. The question was where a client should live, what it should depend on, and how it should report failure.

The decision

It lives in this repository as sdk/elixir, and is zygo_sdk on Hex. The Python and Node clients already have a contract here — a stand-in API, the method table in chapter 17, a test per OpenAPI operation, one version for everything — and a client elsewhere would have to follow it by hand.

It depends on Mint, which is an HTTP connection as a plain value and opens a unix socket, and NimblePool, which lends one of those values to one process at a time. Erlang’s own HTTP client can reach a unix socket, but it keeps its connections to itself, so the rules the other clients test — drop a connection idle for 20 seconds, send again only when nothing came back — could not be kept.

Failures are one exception, Zygo.Error, with a kind such as :busy or :handler. Python needs eleven classes because it branches with except; Elixir branches on data, with case.

  Python:  except zygo.Busy as e:            Elixir:  {:error, %Zygo.Error{kind: :busy}}
           except zygo.HandlerError as e:             {:error, %Zygo.Error{kind: :handler}}

What it costs, and what would reopen it

The Elixir client has two dependencies where the others have none, and Mint brings a third small one. make test-sdk runs three suites; without Elixir the third says it skipped. Publishing needs a Hex key. An :httpc that let a caller hold its own connections would let the client drop Mint. Full ADR 0007.

ADR 0008: A function may listen on its own loopback; a pool may not

The question

From Linux 6.7 Zygo refused every TCP bind in every sandbox, and under network = "none" every connect too, on the reasoning that no mode has ingress so nothing should listen. n8n’s runner launcher could not start in a sandbox: each runner listens on a local health-check port. Neither could anything else that talks to itself over 127.0.0.1 — Jupyter, Ray, PyTorch’s distributed runtime, a headless Chrome. And only on 6.7+; on Debian 12’s 6.1 the same spec worked. The question was what the rule was protecting.

The decision

Ingress is pasta’s job, done on every kernel: it forwards no port into a sandbox, and under none there is no interface but loopback. The bind rule protected one thing — a runtime pool, whose requests belong to different tenants and share a namespace, where a loopback listener is a channel between them. So the rule stays exactly there, and goes everywhere else: a function or a zygo run sandbox may listen on its own loopback, and a sealed function is bounded by its empty namespace rather than by a rule.

  function: one tenant's namespace         pool: shared between tenants
  bind ✓  connect to self ✓                bind ✗ (strict, then Landlock)
  from outside: nothing gets in            from outside: nothing gets in

What it costs, and what would reopen it

Under egress the connect half of talking to yourself still meets the allowlist’s port rule, so a function names the port it listens on. Two tenants’ requests in one namespace outside a pool would need the flag set, not the rule revisited. Landlock gaining an address scope for bind would let pools listen safely too. Full ADR 0008.

ADR 0009: The API may allow private addresses when its operator says so

The question

An embedder’s pool scripts had to call back to the embedder’s own service, on an address in a private range. Chapter 14’s answer for that is an allow rule naming the address and --allow-private-net typed by a person. A spec file served by hand could have it, but nothing served over the API could: the API always sent allow_private_net: false, so that no request body could widen the boundary.

The decision

zygo api --allow-private-net sets it for what deploy callers serve. A body still cannot. The person starting the API types it, as they type --allow-deploy, and it does nothing without that. A rule still names one address and one port. GET /version says private_net.

Two things found on the way are refused now, where they used to be accepted and then fail on every request: a private address written without a prefix, and a network under seccomp = "strict", which has no socket.

What it costs, and what would reopen it

Every deploy caller of such a listener may allow private addresses, not only the one that needed it. Deploy rights are already a shell as the API’s user, so this adds little. A channel for asking the caller mid-run, with no network at all, would reopen it, and so would deploy rights handed to parties the operator does not trust. Full ADR 0009.