Docker-in-Docker Sandbox for container.run
This document is the design record for moving all
container.runstep execution off the host Docker daemon and onto a dedicated Docker-in-Docker (DinD) sandbox service. The DinD sandbox described here is implemented; §2 Current Architecture records the pre-DinD baseline the design started from. The guiding constraint throughout is "everything works as is": existing flows, step parameters, staging conventions, cleanup rules, and the demo lab must keep their observable behavior, one daemon-level down.Status: cutover done. The sandbox is mandatory in every dev and prod stack, not only the demo: the launchers always apply
docker-compose.dind.yml, the worker refuses to start without an attested sandbox, and the host-socket mode is gone. Text below that treats an unsetHEGEMONY_CONTAINER_DOCKER_HOSTas host-socket mode, or as the rollback path, records the transition; §6.5, §7.2 item 6 and §8 are marked superseded where they stand. Operator upgrade steps are in Production Hardening.
Table of Contents
- Overview & Motivation
- Current Architecture
- Target Architecture
- Networking
- Implementation Phases
- Failure Modes & Mitigations
- Testing Strategy
- Rollback
- Behavior-Change Audit
- Decision Record
- Roadmap: Bastion Transport & SOCKS-Aware Probes
- Roadmap: Flow & Step Egress Policies
- Multi-Worker & Remote-Worker Topology
- eBPF Egress Enforcement: Feasibility
1. Overview & Motivation
Today the worker mounts the host's /var/run/docker.sock and runs every container.run step as a sibling container on the host daemon (deploy/compose/docker-compose.dev.yml:183-184, docker-compose.prod.yml:161-162, docs/architecture/overview.md "Container Handler"). Steps that set attach_docker_socket: true hand that same host socket to flow-authored containers — which is full root on the machine.
This design relocates the entire flow-container estate to one privileged docker:dind compose service. Two motivations drive it:
- Blast-radius containment. Flow authors become sandbox admins instead of host root. The host socket disappears from everything a flow can touch:
attach_docker_socketnow grants the DinD daemon's socket,privileged/pid_mode: host/network_mode: hostgrant power over the DinD container only. The API's admin-only gate on these fields (apps/api/routers/flows/lifecycle.py:88-193) stays as defense in depth, but the thing being gated is no longer the host. - One-stroke cleanup. The whole lab/Gitea/step-container/image estate lives inside the DinD service and its data volume. Removing that service and volume clears everything a flow ever created, while API, worker, scheduler, and databases survive untouched. Today the same cleanup means label-scoped sweeps of the shared host daemon (
deploy/compose/Taskfile.yml:659-667) that must carefully avoid other stacks' containers.
The rollout is deliberately reversible: a single new worker setting selects the daemon, and leaving it unset is byte-for-byte today's behavior (see Rollback).
2. Current Architecture
The pre-DinD baseline this design started from, verified against the code at the time; file pointers are the anchors the implementation phases refer back to. Where the implementation has since landed (Phases 1–3), the shipped behavior is what §3 describes.
2.1 Worker ↔ host daemon
- The worker image ships a pinned Docker CLI (client only, no daemon, no buildx/compose plugins), built from the
docker/clisource tag in a builder stage rather than installed fromdownload.docker.com, so its Go stdlib can be patched ahead of upstream's pinned toolchain —deploy/compose/Dockerfile.worker. - The worker service bind-mounts the host socket and joins the docker group via
group_add: ${DOCKER_GID:?...}(docker-compose.dev.yml:175-187,docker-compose.prod.yml:156-165). The demo overlay re-derives the socket GID at runtime withsetprivbecausegroup_adddoes not survive its root→appuser switch (docker-compose.demo.yml:23-38). - All daemon access goes through the docker CLI (
settings.docker_bin, envHEGEMONY_DOCKER_BIN,packages/core/settings.py:106-109). There is no docker SDK usage, and in the pre-DinD baseline there was noDOCKER_HOSTreference anywhere in the repo — every docker subprocess was spawned without anenv=override, inheriting the worker process environment. (The implemented §3.2 plumbing is exactly the introduction of that per-invocationDOCKER_HOSToverride.)
2.2 The container.run handler
The handler ships in the hegemony-steps-container wheel (sibling repo hegemony-step-plugins, built into the worker image at deploy/compose/Dockerfile.worker:88; editable dep at pyproject.toml:213). Its behavior is documented by tests/worker/test_container_handler.py:
build_docker_argsproduces thedocker createargv: labelshegemony.run_id/hegemony.step_run_id/hegemony.handler_id/hegemony.stack, resource limits,-w /workspace,sh -c "set -e\n..."command joining (golden argv attests/worker/test_container_handler.py:645-667). The container namehegemony-<run8>-<step8>-a<attempt>is derived in the handler's execute path (documented atdocs/architecture/overview.mdand thecontainer_cleanup.py:66-68docstring).- Provisioning is
docker cpof stagedtempfile.mkdtemptrees: container roots/attachments,/artifacts(with/artifacts/new),/step_outputs,/workspaceare copied to<container>:/; the run stage isdocker start -aviaasyncio.create_subprocess_execwith line-streamed stdout/stderr; generated artifacts are harvested withdocker cp <container>:/artifacts/new/. <dir>;_preflight_cleanup/_force_remove_containerbracket the lifecycle. attach_docker_socket: trueappends the hard-coded bind-v /var/run/docker.sock:/var/run/docker.sock(tests/worker/test_container_handler.py:474-483).- Shared workspaces: volume mode emits
--mount type=volume,source=<vol>,target=/shared,volume-subpath=<run_id>(Docker Engine 26.1+); bind mode falls back to-v <root>/<run_id>:/sharedand, per the settings docstring, "requires an identical host path under Docker-in-Docker" (packages/core/settings.py:145-165). - The handler's only settings channel is
WorkerHandlerServices.container_runtime()returning aContainerRuntime(docker_bin, shared_workspaces_enabled, shared_workspace_root, shared_workspace_volume, artifact_max_file_size_bytes)(apps/worker/step_handlers/services.py:274-286; the dataclass lives in thehegemony_step_sdkwheel, re-exported atapps/worker/handler_registry.py:24-35).
2.3 Other daemon dependencies
- Worker identity: with
HEGEMONY_WORKER_IDunset, the worker asks the daemon for its own Compose container name —docker inspect -f {{.Name}} $(hostname)— to derive the stable id behind the per-host Temporal queuehegemony-host-<id>(apps/worker/run.py:52-102). This query only makes sense against the daemon that runs the worker container, i.e. the host daemon. - Startup janitor:
cleanup_orphaned_containers()sweepsdocker ps -a --filter label=hegemony.run_id, age-guards byORPHAN_GRACE_SECONDS = 3600, and skips foreignhegemony.stacklabels (apps/worker/container_cleanup.py:59-233, called fromapps/worker/run.py:181).cleanup_stale_shared_workspaces()removes run directories older than 24 hours, but only those whose run the API reports ended (or does not know): a run paused at an approval must keep the plan a latertf.applyreads from/shared. When the API cannot answer it removes nothing named after a run and leaves it to the next start. The price: a run stuck in a non-terminal status (its workflow terminated outside the API) keeps its workspace until it is cancelled or its status is corrected; cancelling it removes the workspace at the next worker start. - Container sweep: while the worker runs, every
HEGEMONY_CONTAINER_SWEEP_INTERVAL_SECONDS(60 by default; 0 turns it off)sweep_containers_of_ended_runs()looks at the step containers on the attested sandbox (same label and stack rules) and asks the API about their runs. A container whose run ended at least two minutes ago (ENDED_RUN_CONTAINER_AGE_SECONDS) belongs to a worker that died: it getsdocker stopwith its own stop timeout, capped at two minutes, and is removed. For a tf step that timeout is its stop grace period, and the tool gets the SIGTERM (seedocs/features/terraform-state.md). Containers of running runs, of runs the API does not know, and of other stacks are left alone. Container-backed steps also have a two-minute heartbeat timeout, so a dead worker's step fails (or is retried) within minutes and its run can end. - Shared workspace lifecycle: the pinned worker creates
<root>/<run_id>on its/data/sharedmount (backed by the named volumehegemony-<env>-shared-workspaces); the run-endcleanup_shared_workspaceactivityrmtrees it (apps/worker/flow_workflow.py:1138-1153,apps/worker/flow_activities.py:1301-1323).
2.4 The demo lab (sibling repo hegemony-demo-data)
Branch claude/control-flow-demo-flows; pointers into that repo.
- "Lab: Provision and tear down demo datacenter" (
src/bundles/30-flows-lab.yaml) is the only flow usingattach_docker_socket: true(all six of itscontainer.runsteps, all on imageghcr.io/srl-labs/clab:0.77.0, allnetwork_mode: host):build_imagerunsdocker buildtwice (FRR router + lab host images).deploy_labrunscontainerlab deploywithprivileged: true,pid_mode: host, andextra_mountsof/var/run/netns:...:shared,/var/lib/docker,/tmp/meridian-lab(30-flows-lab.yaml:91-117).connect_workers("Attach lab network to workers") runsdocker network connect meridian-lab-mgmt <worker>for every container labeledcom.docker.compose.service=worker, grafting the workers onto the clab management bridge so in-process netcli/probe/ shell handlers reach lab devices (30-flows-lab.yaml:123-151).wait_readypollsdocker exec clab-meridian-lab-<node> ....deploy_gitearunsdocker run -d --name meridian-gitea -p 3000:3000 ... gitea/gitea:1.22-rootlessand seeds it (30-flows-lab.yaml:207-250).teardownrunscontainerlab destroy --cleanup, detaches workers, removes themeridian-lab-mgmtnetwork and optionally Gitea (30-flows-lab.yaml:275-333).
- The clab topology pins the management network —
meridian-lab-mgmt,ipv4-subnet: 172.20.30.0/24(src/files/lab/topology.clab.yml:31-33) — with staticmgmt-ipv4per node (topology.clab.yml:48-131); the same IPs appear indemo-inventory/devices/*.yaml(in thehegemony-demo-inventoryrepository) and the ansible inventories. - Gitea addressing is split today: flow steps run host-netted and use
localhost:3000(GITEA_HOST: localhost,src/bundles/05-variables.yaml:26-39; hardcoded clone URL atsrc/bundles/35-flows-ops.yaml:90), while the platform's API-side git integration reaches the host-published port ashttp://host.docker.internal:3000/...withallow_insecure_url: true(src/bundles-meridian-inventory/10-git-repositories.yaml:37-53; composeextra_hostsatdocker-compose.demo.yml:48-51). All git clones/pushes run inside the API process (apps/api/services/git_ops.py; workers never touch git). - The ansible IaC steps (
src/bundles/37-flows-iac.yaml:94-98,725-730) usenetwork_mode: hostto reach172.20.30.x. The terraform flow's containers need no lab reachability (its device push is an in-process netcli step on the worker); it matters here because its plan/apply/ verify steps exercise the per-run/sharedworkspace viaexecution_affinity: shared.
3. Target Architecture
3.1 The dind compose service
A dedicated privileged docker:dind service, one per stack, with:
- a pinned image tag (e.g.
docker:28-dind; the embedded engine must be ≥ 26.1 forvolume-subpathmounts — same floor the host daemon has today,packages/core/settings.py:160-161), - a named data volume at
/var/lib/docker(dind-data). Thedocker:dindimage already declaresVOLUME /var/lib/docker(so an anonymous volume would keep dockerd's overlay2 off the outer overlayfs either way); naming it is what makes the two stories deterministic — persistence across dind recreation (images, containers, inner volumes and networks survive), and one-stroke cleanup (removing this one addressable volume clears the whole flow estate, with no orphaned anonymous volumes left behind), - a fixed IPv4 address on the stack's default compose network. This is the first
networks:/ipam:declaration in the compose layer (no compose file declares any today), so the default network gains an explicit subnet, parameterized per stack to preserve the many-stacks-per-host property (docker-compose.dev.yml:168-170,docker-compose.prod.yml:152-155):
# docker-compose.dind.yml (new overlay — sketch)
services:
dind:
image: docker:28-dind
privileged: true
restart: unless-stopped
environment:
# Empty disables dockerd's TLS auto-provisioning so it serves
# plain TCP on 2375 — acceptable only because the port is never
# published and the compose network is stack-internal (§6.3).
- DOCKER_TLS_CERTDIR=
command:
# No --host flags: dockerd-entrypoint.sh already serves
# unix:///var/run/docker.sock plus tcp://0.0.0.0:2375 when
# DOCKER_TLS_CERTDIR is empty, and repeating them aborts dockerd
# with duplicate listeners. Only extra daemon options belong here.
- --mtu=${HEGEMONY_DIND_MTU:-1500}
networks:
default:
ipv4_address: ${HEGEMONY_DIND_IPV4:-172.28.100.10}
volumes:
- dind-data:/var/lib/docker
- shared-workspaces:/data/shared
ports:
# Optional: host-browser access to the demo Gitea UI. The inner
# `docker run -p 3000:3000` binds :3000 in the dind netns; this
# publish re-exports it to the host.
- "${HEGEMONY_DIND_GITEA_BIND:-127.0.0.1}:3000:3000"
healthcheck:
test: ["CMD", "docker", "info"]
interval: 10s
timeout: 5s
retries: 12
start_period: 30s
networks:
default:
ipam:
config:
- subnet: ${HEGEMONY_COMPOSE_SUBNET:-172.28.100.0/24}
volumes:
dind-data:Subnet/IP defaults differ per env file so a dev, prod, and demo stack can coexist on one host (.env.dev / .env.prod / .env.demo each pin their own HEGEMONY_COMPOSE_SUBNET + HEGEMONY_DIND_IPV4). The chosen ranges must avoid the lab subnet 172.20.30.0/24 and the inner dind default bridge (172.17.0.0/16 from dockerd defaults).
The compose network is dual-stack: enable_ipv6: true plus a ULA v6 subnet (HEGEMONY_COMPOSE_SUBNET_V6), so every stack container gets an IPv6 address that Docker masquerades to the host's v6 uplink. The dind daemon runs --ipv6 --ip6tables with its own inner default-bridge ULA (HEGEMONY_DIND_FIXED_CIDR_V6), so flow containers and the egress enforcer's dual-stack policy networks (§12.4) reach external IPv6 via a double NAT66: inner ULA → dind's compose-network v6 (HEGEMONY_DIND_IPV6) → the host. Each v6 ULA is distinct per stack, the same way the v4 subnets are. A host without an IPv6 uplink can leave the v6 vars unset in the base env and drop --ipv6 to run v4-only.
3.2 Daemon selection: HEGEMONY_CONTAINER_DOCKER_HOST
One new worker setting selects the daemon for flow-container work only:
packages/core/settings.pygainscontainer_docker_host: str = Field(default="")(envHEGEMONY_CONTAINER_DOCKER_HOST), next todocker_binin the existing "Container Execution (DinD)" section (settings.py:102-121). Unset ⇒ current host-socket behavior, which is also the rollback story.- The handler receives it through the existing settings channel:
ContainerRuntime(in thehegemony_step_sdkwheel) gains adocker_host: str = ""field, populated byWorkerHandlerServices.container_runtime()(apps/worker/step_handlers/services.py:274-286). - The handler passes it as
DOCKER_HOSTper invocation — everysubprocess.run/asyncio.create_subprocess_execdocker call getsenv={**os.environ, "DOCKER_HOST": docker_host}when the setting is non-empty. Thedocker createargv frombuild_docker_argsis untouched (daemon selection is environment, not argv — the golden argv test attests/worker/test_container_handler.py:645-667must keep passing verbatim). apps/worker/container_cleanup.pyapplies the same env override to itsdocker ps/docker inspect/docker rmcalls (container_cleanup.py:91-105,136-149,197-204), so the janitor sweeps the same daemon the handler creates containers on.
Deliberately not redirected: the worker-id derivation (apps/worker/run.py:66-72). It asks "what is the name of the container whose id is my hostname" — a question only the daemon hosting the worker can answer. This is why DOCKER_HOST must never be exported process-wide: it would silently break id derivation and strand pinned runs on a dead per-host queue. The worker therefore keeps its host-socket mount (and DOCKER_GID); retiring it is tracked in TODO.md. Note the containment goal is still met: the host socket stays a worker-level privilege — no flow-authored container can reach it anymore.
3.3 Wholesale relocation and parameter semantics
Everything a flow creates moves under the one DinD daemon: the clab runner, lab nodes, Gitea, ansible/terraform step containers, and images built by flows. Existing step parameters keep their exact semantics one level down, because bind sources, network names, and namespaces all resolve against the daemon that runs the container:
| Step param | Today (host daemon) | Under DinD |
|---|---|---|
network_mode: host | host netns | the dind container's netns |
network_mode: <name> | host daemon's network | dind daemon's network (clab creates meridian-lab-mgmt there) |
privileged, pid_mode: host | host-scoped | dind-scoped |
extra_mounts: /var/run/netns:...:shared | host's netns dir | dind's netns dir (where clab's node netns now live) |
extra_mounts: /var/lib/docker | host daemon state | dind daemon state (the dind-data volume) |
extra_mounts: /tmp/meridian-lab | host /tmp | dind container /tmp (still shared across the lab flow's steps) |
attach_docker_socket: true | host socket | dind's own socket (§3.4) |
localhost:3000 from a host-net step | host-published Gitea | Gitea's publish binding in the dind netns — still works |
3.4 attach_docker_socket resolution
No handler change is needed for the mount itself. The handler's hard-coded -v /var/run/docker.sock:/var/run/docker.sock (tests/worker/test_container_handler.py:474-483) is a bind whose source resolves against the filesystem of the daemon executing the create — under DOCKER_HOST=dind that is the dind container's root filesystem, where /var/run/docker.sock is the dind daemon's own unix socket (the entrypoint keeps serving it alongside TCP — see §3.1). Socket-attached steps therefore transparently become clients of the sandbox daemon.
A shared socket volume (dind mounting a named volume over /var/run, the worker mounting it read-only elsewhere) is not required for step attachment; it is only the transport for the hardened worker→dind unix-socket alternative discussed in §6.3.
3.5 Per-run /shared workspaces on the dind daemon
Volume mode stays the mechanism, so the handler's --mount type=volume,source=<vol>,target=/shared,volume-subpath=<run_id> argv is unchanged — but the named volume must now exist on the dind daemon, and it must be the same bytes the worker sees at /data/shared (the worker creates <root>/<run_id> before pinned steps and rmtrees it at run end; docker cp staging is daemon-agnostic and needs nothing).
Mechanism: mount the existing compose volume into the dind service at /data/shared (§3.1 sketch), then register a bind-backed named volume on the dind daemon pointing at that path:
docker volume create --driver local \
--opt type=none --opt o=bind --opt device=/data/shared \
hegemony-dev-shared-workspaces # = HEGEMONY_SHARED_WORKSPACE_VOLUMEvolume-subpath=<run_id> then resolves inside the same directory tree the worker manages, so workspace creation, the run-end cleanup activity, and the stale-workspace janitor (container_cleanup.py:258-305) all keep working without modification. The docker volume create is idempotent and is executed by a small dind-init one-shot service in the overlay (docker:28-cli image, depends_on: dind: service_healthy, DOCKER_HOST=tcp://dind:2375). The volume record lives in dind-data, so it survives dind restarts; recreating the stack re-runs the init.
3.6 API ↔ Gitea git integration
The demo git_repositories clone URLs and the GITEA_URL demo variable change to http://dind:3000:
- The API resolves
dindvia compose DNS (same default network) and its git client accepts the URL because the repositories already carryallow_insecure_url: true— the SSRF/private-address guard is skipped exactly for this demo-Gitea case (apps/api/services/git_ops.py:86-138,tests/git/test_git_ops.py:50-60). - Host-netted flow steps reach the same
http://dind:3000: they share the dind netns, where compose DNS (127.0.0.11) resolvesdindand the inner-p 3000:3000publish is bound. One canonical URL replaces today'slocalhost:3000/host.docker.internal:3000split. - Optional host-browser access to the Gitea UI comes from the dind service's own
ports:publish (§3.1), keepinghttp://localhost:3000working from the operator's machine. - The
extra_hosts: host.docker.internal:host-gatewayentries (docker-compose.demo.yml:48-51,63-64) stop being needed for Gitea; they stay for any other host-published integrations.
4. Networking
Superseded (demo data path): the routed worker→lab mechanism this section designed (§4.1–§4.3) shipped, then was deleted by the §11 B4 bastion cutover — no host route, no
lab-route-installersidecar, no dind-netnssysctl/iptableslegs exist anymore. Workers reach lab devices only through the bastion's authenticated listeners (dind:2222SSH ProxyJump,dind:1080SOCKS5). §4.1–§4.3 are kept as the design record of the interim mechanism; §4.4's MTU guidance still applies to the dind service's own networks.
4.1 Worker → lab data path
In-process netcli/probe/shell handlers run inside the worker and must reach lab devices at 172.20.30.x. Today the demo grafts the workers onto the clab bridge with docker network connect. Under DinD the lab bridge exists inside the dind daemon, and the chosen mechanism (over netns-sharing alternatives) is routing: a host-level route pointing the lab subnet at dind's fixed compose IP — mirroring the 172.20.30.0/24 dev br-<clab> route Docker itself installs on the host today — plus forwarding and firewall openings inside dind.
Packet walk (worker 172.28.100.x → device 172.20.30.11):
- The worker has no route for
172.20.30.0/24, so the packet goes to its default gateway — the host's address on the compose bridge. - The host routing table forwards it back onto the same bridge toward dind (rule below). Docker's own
-i br-X -o br-X -j ACCEPTFORWARD rule permits same-bridge routed traffic; the host may emit an ICMP redirect telling the worker to use dind directly next time (harmless either way). - dind forwards it from
eth0onto the inner clab bridge. dockerd enablesnet.ipv4.ip_forwarditself, but the inner FORWARD policy is DROP, so explicit accepts are needed (rules below). - Replies are conntrack-ESTABLISHED and are neither filtered nor masqueraded on the way back (MASQUERADE only rewrites NEW outbound flows). Lab-initiated traffic to the platform (e.g. device syslog) leaves dind masqueraded as the dind IP — acceptable for the demo.
4.2 Exact rules
Host (root netns) — one static route:
ip route replace 172.20.30.0/24 via ${HEGEMONY_DIND_IPV4}
# verify: ip route get 172.20.30.11Inside the dind netns — forwarding + firewall openings (idempotent):
sysctl -w net.ipv4.ip_forward=1 # dockerd default, asserted for determinism
iptables -C DOCKER-USER -d 172.20.30.0/24 -j ACCEPT 2>/dev/null \
|| iptables -I DOCKER-USER -d 172.20.30.0/24 -j ACCEPT
iptables -C DOCKER-USER -s 172.20.30.0/24 -j ACCEPT 2>/dev/null \
|| iptables -I DOCKER-USER -s 172.20.30.0/24 -j ACCEPTDOCKER-USER is evaluated before Docker's own FORWARD rules and is the supported place for operator rules; the accepts cover both directions, including the ICMP needed for path-MTU discovery.
4.3 Who installs what
- dind-internal rules (§4.2 second block): the demo's "Attach lab network to workers" flow step is repurposed to install them. Relocated to the dind daemon, its
network_mode: hostnow means the dind netns — exactly where these rules live — and it gainsprivileged: trueforsysctl/iptables(the lab flow already uses privileged host-netns containers fordeploy_lab/teardown). Its olddocker network connect <worker>loop is dropped: the workers are not containers of the dind daemon, and the label filter would match nothing there. Because dind-netns rules evaporate on a dind restart, the step stays a normal part of every lab deploy (and is safe to re-run any time). - The host route (§4.2 first block): a dind-hosted step cannot write the host routing table — under the sandbox model nothing a flow runs can, which is the point of the exercise. This forces one amendment to the decided sketch (which had the flow step install all three legs): the route leg moves to stack provisioning. It is one static, lab-independent line (subnet and next-hop are stack constants), installed by a minimal sidecar in the demo overlay (
network_mode: host,cap_add: [NET_ADMIN],ip route replace ... && sleep infinity,restart: unless-stoppedso it re-asserts after host reboots), or equivalently a documentedinstall.sh/ manual one-liner. The amendment and its eventual deletion by the bastion cutover are recorded in §10.
No compose-level netns sharing, no macvlans, and no worker-service changes are involved.
4.4 MTU nesting
veth/bridge hops add no encapsulation, so nesting does not shrink MTU by itself. The risk is an outer path MTU below 1500 (VPN/overlay uplinks, some cloud fabrics): Docker bridges default to 1500 regardless, and under DinD there are now three places that must agree —
- the compose default network (host daemon),
- dind's own networks (
--mtuflag in §3.1, exposed asHEGEMONY_DIND_MTU), - the clab management network (topology-level
mtuif ever needed; the demo topology sets none —src/files/lab/topology.clab.yml).
Rule of thumb: inner MTU ≤ outer effective MTU. Defaults (all 1500) are correct for plain Ethernet hosts; on constrained uplinks set HEGEMONY_DIND_MTU to the compose network's MTU. (With the §11 bastion data path, worker→lab traffic rides TCP tunnels, so path-MTU issues surface as ordinary in-tunnel TCP behavior rather than blackholed PMTUD.) Verification: docker network inspect / ip link inside dind.
5. Implementation Phases
Each phase is independently shippable and leaves behavior unchanged until the final env flip. Phases 1–2 span two repos because the handler ships in the hegemony-step-plugins wheel.
Phase 1 — SDK + handler plumbing (hegemony-step-plugins repo)
hegemony_step_sdk: adddocker_host: str = ""toContainerRuntime.plugins/steps_container(hegemony_steps_container.run): threaddocker_hostinto every docker subprocess as a per-invocationenv={**os.environ, "DOCKER_HOST": ...}—docker create, alldocker cpstaging/harvest calls,docker start -a(asyncio.create_subprocess_exec),_preflight_cleanup,_force_remove_container. No argv changes;attach_docker_socketkeeps its hard-coded mount (§3.4).- Unit tests in that repo for env construction (set vs unset).
- Release the wheels.
Phase 2 — Platform plumbing (this repo)
packages/core/settings.py(afterdocker_bin, ~line 109): newcontainer_docker_hostfield with a docstring covering the unset-means-host-socket contract and the per-invocation rule.apps/worker/step_handlers/services.py:274-286: passdocker_host=settings.container_docker_hostintoContainerRuntime.apps/worker/container_cleanup.py: build the env override once incleanup_orphaned_containers()and passenv=to the three subprocess call sites (:91-105,:136-149,:197-204).apps/worker/run.py:52-80: comment stating worker-id derivation intentionally targets the daemon hosting the worker and must not honorcontainer_docker_host.- Bump the step-plugins pin (
pyproject.toml:71,213); worker image rebuild picks up the new plugin source (deploy/compose/Dockerfile.worker:88). - Test extensions per §7.1.
Phase 3 — Compose layer (this repo)
- New overlay
deploy/compose/docker-compose.dind.ymlper the §3.1 sketch:dindservice (privileged, pinned image,dind-data+shared-workspacesmounts, fixed IP, healthcheck, optional:3000publish), the ipam'ddefaultnetwork, thedind-initvolume-create one-shot (§3.5), and the worker additions (HEGEMONY_CONTAINER_DOCKER_HOST=tcp://dind:2375,depends_on: dind: service_healthy). - Register the overlay:
deploy/compose/Taskfile.yml:81(VALID_OVERLAYS) anddeploy/compose/dc.sh:47-54; include it in the default service lists (Taskfile.yml:31-36) only at cutover. (EXTRA_FILESindc.sh:63-66allows unregistered experimentation meanwhile.) - Env examples:
HEGEMONY_COMPOSE_SUBNET,HEGEMONY_DIND_IPV4,HEGEMONY_DIND_MTU,HEGEMONY_CONTAINER_DOCKER_HOSTin.env.dev.example/.env.prod.example/.env.demo, with distinct per-stack subnets. deploy/compose/Taskfile.yml:659-667(demo:reset): the host-daemon label sweep keeps running (legacy/host-mode containers) but the primary cleanup becomes the stack'sdown --volumesremovingdind-data; add a dind-side sweep (docker exec <dind> sh -c 'docker ps -aq --filter label=hegemony.stack=... | xargs -r docker rm -f') for resets that keep volumes.- Demo overlay: the host-route sidecar (§4.3), parameterized by
HEGEMONY_LAB_ROUTE_SUBNET(value lives in.env.demo) — since deleted by the §11 B4 bastion cutover. security:trivy:iac(rootTaskfile.yml:372-377) scansdeploy/composewithtrivy config; compose files are not currently a Trivy misconfig target, so the privileged service likely passes — verify on the branch and add a scoped, justified ignore only if a finding appears.- Docs: update
docs/architecture/overview.md:286-327(container handler),docs/deployment/production-hardening.md:46-70(the socket-exposure section gets its primary mitigation),docs/demo.mdtroubleshooting (:208), and the compose comments that currently say "sibling step container". - Watch the string-pinning deploy tests (
tests/deploy/test_demo_bootstrap_orchestration.py:37-146) when touchingdocker-compose.demo.yml/Taskfile.yml/.env.demo.
Phase 4 — Demo-data changes (hegemony-demo-data repo)
These land after Phases 1–3 are released (the bundles assume the new daemon), and require a demo reset to take effect — bootstrap import is one-shot per database (docs/demo.md, "reset to re-run instance bootstrap").
src/bundles/05-variables.yaml:26-39:GITEA_HOST: localhost→dind(GITEA_URLderives; description text updated).src/bundles/30-flows-lab.yaml:deploy_gitea(:207-250):GITEA__server__ROOT_URLfrom{{ vars.GITEA_URL }}instead of the hardcodedhttp://localhost:3000/; consider--restart unless-stoppedso Gitea returns after a dind restart.- approval-gate message (
:269-274): browser URL note (host access now via the dind service publish). connect_workers(:123-151): repurposed per §4.3 — rename to reflect "open lab route", addprivileged: true, replace thedocker network connectloop with the idempotent sysctl/iptables block.teardown(:275-333): drop the worker-detach loop (:316-317) — nothing attaches workers anymore; the rest (containerlab destroy, network rm, Gitea rm) works unchanged against dind.
src/bundles/35-flows-ops.yaml:90: clone URLhttp://localhost:3000/...→{{ vars.GITEA_URL }}/....src/bundles-meridian-inventory/10-git-repositories.yaml:37-53:http://host.docker.internal:3000/...→http://dind:3000/...(keepallow_insecure_url: trueand the auth refs).src/files/gitea/seed.sh:31: update the fallback default to match.- Regenerate
dist/(hegemony-demo-data/scripts/build.py), updateREADME.md/hegemony-demo-data/docs/walkthrough.mdprose aboutlocalhost:3000andhost.docker.internal.
Sequencing within the phase: bundle changes are atomic per reset — there is no mixed state to support. Sequencing across repos: a platform stack still in host-socket mode with the new bundles would break (the dind hostname resolves, but Gitea would deploy on the host daemon while URLs point into the compose network), so the demo-data release notes must pin the minimum platform version, matching the existing wheel-pinning practice (hegemony-demo-data/deploy/compose/demo-plugin-wheels.txt).
Phase 5 — Cutover & follow-ups
Status: done. The launchers apply
docker-compose.dind.ymlto every stack,SERVICES=coreincluded, anddindsurvives only as a deprecated no-op name; compose fills intcp://dind:2375for an unset or emptyHEGEMONY_CONTAINER_DOCKER_HOST, and the worker refuses to start on an empty endpoint or token, or on a daemon that answers but is not the attested sandbox. The host sweep is an operator step, documented with the upgrade in Production Hardening rather than automated. The list below is the plan as it stood.
The non-demo cutover was outstanding — tracked in TODO.md:
- Flip the default service lists to include the
dindoverlay; setHEGEMONY_CONTAINER_DOCKER_HOSTin the shipped env examples. - One-time hygiene at cutover: sweep leftover flow containers from the host daemon (
docker ps -a --filter label=hegemony.run_id) — the janitor now watches the dind daemon and will not see them. - Follow-ups — retiring the worker's host socket and resource limits on the dind service — are also tracked in
TODO.md.
6. Failure Modes & Mitigations
6.1 Cold image cache on first boot
A fresh dind-data volume knows no images: the first lab run pulls ghcr.io/srl-labs/clab, gitea/gitea, alpine, hashicorp/terraform, and docker build pulls quay.io/frrouting/frr — all through the dind daemon's egress (compose network NAT). First-run step latency rises accordingly, and build_image's timeout_seconds: 600 now also covers the base-image pulls.
The egress helper image is now pre-warmed by dind-init, which pulls container_egress_helper_image right after registering the volumes. It is the one image whose cold pull is not attributable to anything the user asked for: a policied step's helper docker run fetched a few hundred MB implicitly, stalling the step ~18s before it could emit its first event. The pull is deliberately best-effort — the worker gates on this one-shot completing (service_completed_successfully), so a registry hiccup must degrade to the old inline pull rather than stop the stack from booting. The compose expression matches the setting's default and resolves from the same --env-file the worker's env_file supplies, so the two sides cannot name different images.
Remaining mitigations for the rest: the data volume makes this a once-per-volume cost; the same dind-init extension can front-load further images if a deployment wants them; step timeouts in the demo bundles get headroom reviewed during the Phase 4 smoke run; a registry mirror (--registry-mirror on the dind dockerd) is available for constrained networks (a pull-through cache is tracked in TODO.md).
Flow-step images are deliberately not pre-warmed: which ones a stack needs is a property of its flows, not of the platform, and the container.run pull now reports byte progress as it goes, so the wait is attributable rather than mysterious.
6.2 dind restart semantics
A dind restart (crash, docker compose restart dind, host reboot) stops every inner container. Images, inner volumes, inner networks, and the volume registration from §3.5 persist in dind-data; what is lost:
- Running lab + Gitea: not restarted automatically (clab nodes and the demo's
docker run -dcarry no restart policy). Recovery is the lab flow itself —containerlab deploy --reconfigureis idempotent by design, and Gitea's data loss is already accepted demo behavior (docker rm -fon redeploy today). Optionally--restart unless-stoppedon Gitea (Phase 4). - Worker→lab reachability: nothing to re-install. The §11 B4 cutover deleted the routed path entirely — no dind-netns route legs, no host route, no sidecar. Workers reach lab devices through the bastion the lab flow deploys, so reachability returns with the bastion container itself;
containerlab deploy --reconfigurerestores it. - Steps mid-flight: their
docker start -aclient fails; the activity fails and Temporal retries per step policy. Orphaned containers on dind are reclaimed by the startup janitor on the next worker restart (same 1-hour grace,container_cleanup.py:23), by the dind-side reset sweep, or, once their run has ended, by the periodic container sweep (below). After a host reboot that reclamation is now automatic: the worker carriesrestart: unless-stoppedtoo, so it comes back and runs its startup janitor without an operator.
restart: unless-stopped on the dind service bounds the outage window, and every other long-running service in the stack carries it as well — previously dind came back at boot into a stack that did not, so it had no clients.
Two boot-order hazards are closed explicitly, because daemon auto-start honours no compose ordering:
- Fixed-address theft. dind is the stack's only fixed-IP container, and boot-time IPAM hands out dynamic addresses sequentially from the bottom of the subnet — a sibling could draw dind's address first, turning dind's start into a permanent "Address already in use" failure no restart policy retries out of. The compose network's
ip_rangenow confines dynamic allocation to the upper part of the subnet; fixed addresses live below it (tests/deploy/test_dind_overlay.pypins the invariant for the defaults and every env file). - Worker-before-dind. The worker used to poll Temporal within seconds of boot while dind was still starting, so its startup janitors silently skipped their sweeps and container.run steps dispatched in the window failed on daemon connection errors.
wait_for_sandbox_daemon(apps/worker/run.py) now holds janitors and pollers back until the sandbox answers, bounded byHEGEMONY_CONTAINER_DOCKER_WAIT_SECONDSso a genuinely dead sandbox degrades container steps instead of all work. With the wait at0the worker does not block but still probes once, so a daemon that answers is attested, or refused, before the worker connects to Temporal. Past the deadline, or when that one probe fails, the worker starts without an attested sandbox:attest_sandbox_when_reachablekeeps probing at the same interval and logs why it has not attested yet, andWorkerHandlerServices.container_runtime()refuses every container step, quoting that reason, until that probe's daemon has passed the §12.2 marker check (the process-wideSANDBOX_ATTESTATIONinegress_enforcement.py). The worker's heartbeat advertises thecontainerscapability only once that check has passed, so shared-worker assignment pins no newsharedorexplicitrun to a worker that would refuse its container steps. A worker keeps its id across restarts, so startup first sends one heartbeat with no capabilities, before the bounded wait, to clear what the previous process advertised; assignment always asks forcontainers, so that heartbeat does not make the worker eligible before its host queue has a consumer. Until the check passes the heartbeat goes out with no capabilities, which keeps the worker online, and the attestation wakes it when it ends, so the capability is advertised at once rather than a heartbeat interval later. A run pinned to the worker before a restart keeps its pin (assignment returns an existing pin as it is), so its container steps are refused until the sandbox is attested. The docker janitors sweep only a daemon that has passed: at startup, or after a degraded start between the marker check and enabling container steps, so a mispointed endpoint is never swept. A late daemon that fails the check stops the worker with a non-zero exit, exactly as it would have been refused at startup; one whose marker cannot be read at all (it stopped answering in between) is probed again instead.
dind-init deliberately stays restart: "no": the daemon never auto-starts a one-shot, and it does not need to, because everything it registers (the bind-backed shared-workspace volume, the hegemony-sandbox-marker label, the pre-pulled egress helper image) lives in dind-data and survives. The residual edge is losing dind-data without recreating the stack — dind-init would stay Exited (0), the volume registration would be gone, and the first container.run step's --mount would auto-create the empty volume this design warns about. In practice dind-data only goes away with down -v, which removes dind-init too and so forces it to re-run on the next up.
6.3 TLS-off TCP vs mounted-socket transport
The worker→dind control channel has two viable shapes:
tcp://dind:2375 (TLS off) | shared unix-socket volume | |
|---|---|---|
| Exposure | Unauthenticated daemon API on the compose network (and reachable from inner step containers). Never published to the host. | No network listener; socket file permissions gate access. |
| Setup | DOCKER_TLS_CERTDIR= only (the entrypoint opens both listeners). | Shared volume over dind's /var/run + worker mount + gid alignment (details below). |
| Failure surface | Network-level (conn refused ⇒ clear step errors). | Volume/permissions-level (EACCES until gids align). |
Socket-volume setup, in full: dind mounts a named volume over /var/run; the worker mounts the same volume (e.g. at /var/run/dind) and sets HEGEMONY_CONTAINER_DOCKER_HOST=unix:///var/run/dind/docker.sock; the dind dockerd runs with --group aligned to the worker's gid. Stale socket files across dind restarts are harmless (dockerd re-binds).
The recommendation is TCP on the internal network as the default: it is operationally simpler, and the exposure it adds is within the trust boundary this design defines — every flow author is already a sandbox admin (any attach_docker_socket step gets full dind control by design), and platform services on the compose network are trusted. The hard rules are: never publish 2375/2376, and never point the setting at anything but the stack's own dind. The socket-volume variant remains the documented hardening option for deployments that want zero unauthenticated listeners; the setting accepts either URL form unchanged. TLS-on TCP (2376 with cert provisioning) is intentionally out of scope — it buys little inside a single-host stack and adds a cert lifecycle.
6.4 Disk growth inside dind-data
Images, stopped containers, build cache, and inner volumes accumulate in dind-data with no GC of their own.
Partly resolved by the sandbox_image_pruning maintenance job (apps/api/services/maintenance/jobs/sandbox_image_pruning.py), which runs on the ordinary maintenance cadence (6h by default) and sweeps the two regenerable categories over the daemon's Engine API: dangling images and build cache. It attests the daemon with the §12.2 marker token before deleting anything — a prune aimed at a mispointed endpoint would delete an operator's own images. A missing endpoint or token, an unreachable daemon, and a daemon that answers without the marker or with another token all fail the job run, so a sandbox that is never pruned shows in platform health instead of as a clean skip. The dind overlay gives the api service the endpoint and token for exactly this reason.
Deliberately still uncovered:
- Stopped containers belong to
container_cleanup.py, which applies a grace period the API cannot see; sweeping them from the prune job would race a step another replica is provisioning. - Tagged-but-unused images are the remaining growth term, and age filters do not help:
docker image prune --filter until=matches an image's build date, not when it was last pulled or used, so it evicts a constantly-usedalpine:3.22while sparing a lab image built this morning. Reclaiming these wants a last-used signal the daemon does not expose, so it stays operator-driven (docker system prune -ainside dind, ordown+ volume rm). - Inner volumes can hold flow data and are never swept automatically.
A size cap on dind-data is the other half of the problem (tracked in TODO.md with the dind resource limits): pruning slows the fill rate, it does not bound it.
6.5 dind unavailable / not ready
Superseded (cutover): worker startup now does require the sandbox to be configured — an empty endpoint or token, or a daemon that answers without the attested marker, makes the worker exit non-zero. What still holds is the degraded start below for a sandbox that is configured but not answering: the worker waits a bounded time, then starts, refuses every container step (a refused step fails without a retry) and advertises no
containerscapability to shared-worker assignment until it has attested the sandbox in the background, and exits non-zero if the daemon that finally answers fails attestation (§6.2).
Worker startup does not require dind: the janitor's failure handling already degrades to warnings (container_cleanup.py:107-109,230-233). A container.run step against a down dind fails its docker create with a clear connection error and follows normal step-failure/retry policy. The overlay's depends_on: service_healthy ordering avoids the window at stack bring-up.
6.6 Version and feature skew
The worker's pinned CLI (v29, Dockerfile.worker:21) negotiates API versions with the dind engine; pin the dind image tag and record the volume-subpath floor (engine ≥ 26.1). Flow-side docker CLIs (the clab image's) already talk to a same-generation daemon today. The worker image still lacks buildx/compose plugins — irrelevant here, since flows do their builds inside step containers, not through the worker's CLI.
7. Testing Strategy
7.1 Unit-testable (fast, no daemon)
- Argv invariance: the
build_docker_argsgolden test (tests/worker/test_container_handler.py:645-667) must pass unchanged — daemon selection must not leak into argv. - Env construction (both repos): with
ContainerRuntime.docker_hostset, every mocked subprocess seam (subprocess.runfakes, theasyncio.create_subprocess_execmock) receivesenv=containingDOCKER_HOST=<value>alongside inherited vars; with it unset, no override is passed. The existing fakes already accept**kwargs(test_container_handler.py:831-851), so assertions bolt on. - Janitor plumbing: extend
tests/worker/test_container_cleanup.py(settings fake at:137-150) to assert the env override ondocker ps/inspect/rm, and its absence when unset. - Settings: default-empty round-trip for
HEGEMONY_CONTAINER_DOCKER_HOST. - Worker-id isolation: a regression test that
run.py'sdocker inspectis invoked without the override even when the setting is set. - Compose assertions: extend the raw-YAML string-pinning house pattern (demo tests) to the new overlay — dind service present, privileged,
ipv4_address/ subnet variables wired, worker env set,dind-initvolume name equal toHEGEMONY_SHARED_WORKSPACE_VOLUME. (Adocker compose configrendered-merge check would be stronger but adds the compose CLI as a test dependency — optional.)
7.2 Live-lab smoke (manual/nightly; needs a privileged daemon)
Not CI-able on standard runners; run on a dind-enabled dev or demo stack:
- Basic step: a plain
container.run(alpine echo) lands on dind (docker exec <dind> docker psshows the labeled container), stdout streams, artifacts harvest. - Full demo lab flow: build → deploy → bastion deploy →
wait_ready→ Gitea deploy + seed → backup flow pushes tohttp://dind:3000→ API-side git sync (test-connection, flow sync) → terraform flow (exercisesvolume-subpathon the dind-side volume) → teardown. - Reachability: netcli/shell from the worker to
172.20.30.xthrough the bastion (ProxyJump viadind:2222);tcp_connectprobe withsocks_proxyviadind:1080; both refused with wrong bastion credentials. - Restart drill:
docker compose restart dindmid-lab; verify janitor cleanup, lab redeploy idempotence, rules re-installation. - Cold-cache timing: fresh
dind-data, measure first lab run against step timeouts. - Rollback drill: unset the setting, restart the worker, verify a step lands on the host daemon again. Superseded (cutover): an unset or empty setting now resolves to the sandbox and the worker refuses any other daemon, so there is no host-daemon state to drill; rollback is redeploying the previous release (§8).
8. Rollback
Superseded (cutover): there is no setting-level rollback any more. Compose resolves an unset or empty
HEGEMONY_CONTAINER_DOCKER_HOSTto the sandbox, the worker refuses to start against a daemon that is not the attested sandbox, and each step refuses an empty endpoint, so the host socket the worker still mounts serves worker-id derivation only. Rolling back means redeploying the previous release. The text below is the transition-era plan, kept as history.
The rollback story is the setting itself:
- Unset
HEGEMONY_CONTAINER_DOCKER_HOST(or start without the overlay) and restart workers: every docker invocation inherits the plain environment again and hits the host socket, which is mounted throughout the transition. No code paths are removed; nothing else needs reverting. - The dind estate is inert after rollback: leave it (its containers age into janitor-invisible stillness — the janitor follows the setting), or clear it wholesale (
down+dind-dataremoval). - Demo-data is the one coupled piece: bundles from Phase 4 assume the dind daemon (URLs, the bastion hop), so rolling the platform back on a demo stack means restoring the previous bundle release and resetting the demo DB — same minimum-version discipline as the plugin wheels, in the opposite direction.
9. Behavior-Change Audit
Everything a flow author or operator can observe, and how the plan preserves (or consciously changes) it:
| # | Observable | Today | Under DinD | Preservation |
|---|---|---|---|---|
| 1 | Step params and staging paths (§3.3 table) | host-daemon semantics | identical, one level down | unchanged by construction — argv identical |
| 2 | attach_docker_socket grants | host root | dind admin | intended change (the motivation); API admin gate unchanged |
| 3 | docker ps on the host shows flow containers | yes | no — inspect via docker exec <dind> docker ps or DOCKER_HOST | docs + Taskfile helpers (Phase 3) |
| 4 | Lab/Gitea survive compose down | yes (host-daemon residents) | stopped with the stack; down -v erases them | intended change (§1); redeploy is idempotent |
| 5 | Gitea from the host browser at localhost:3000 | via flow's -p 3000:3000 | via the dind service's optional publish | preserved when the publish is enabled (default in demo env) |
| 6 | Gitea from API: host.docker.internal:3000 | via extra_hosts | http://dind:3000 over compose DNS | Phase 4 URL change; allow_insecure_url already set |
| 7 | Gitea from host-netted steps: localhost:3000 | host netns port | still works (dind netns port); canonical URL becomes dind:3000 | Phase 4 unifies on GITEA_URL |
| 8 | Worker→lab SSH source address | worker's grafted 172.20.30.x address | the bastion's lab-side address (§11) | lab sshd accepts any source; no device config depends on it |
| 9 | "Attach lab network to workers" step | docker network connect per worker | removed — the lab flow deploys a bastion instead (§11 B4) | same flow position; reachability is now an authenticated hop, not a route |
| 10 | task compose:demo:reset sweep | host-daemon label sweep | down -v removes the estate; dind-side sweep added; host sweep kept for legacy | Phase 3 |
| 11 | First-run latency | host image cache warm | cold dind cache once per volume | §6.1 mitigations |
| 12 | Flow-built images (meridian/*) land on | host daemon | dind daemon | intended; host daemon stays clean |
| 13 | /tmp/meridian-lab shared workdir | host /tmp | dind container /tmp | still shared across the flow's steps; wiped with the service — teardown already tolerates absence |
| 14 | Worker id / pinned-run queues | derived via host socket | unchanged — derivation deliberately not redirected (§3.2) | regression-tested (§7.1) |
| 15 | Janitor scope | host daemon, stack-scoped | dind daemon, stack-scoped (same rules) | env plumbing only; grace/label logic untouched |
| 16 | DOCKER_GID requirement | required | still required (host socket retained for §3.2) | retirement tracked in TODO.md |
10. Decision Record
Questions the design phase raised and settled, kept as history. Outstanding follow-ups — retiring the worker's host socket, dind resource limits, a registry pull-through cache, SSH host-key pinning, and the transport/publish defaults — are tracked in TODO.md.
Adoption. Mandatory in dev and prod, with no host-socket escape hatch.
docker-compose.dind.ymlstays a file of its own, which every launcher applies, rather than being folded into the two base files: one definition serves the launchers and the appliance, which vendors the file. The worker fails closed when the sandbox is not configured or not attested, and degrades rather than exits when it is merely late: it refuses container steps, and withholds thecontainerscapability from shared-worker assignment, until a background attestation passes, and exits if that attestation fails. Each environment pins its own subnet (dev172.28.100.0/24, prod172.28.101.0/24, demo172.28.102.0/24). Leftover flow containers on the host daemon are a documented operator sweep rather than an automatic one, because an automatic sweep would delete on the host without the operator's involvement.Host-route installer. The decided sketch had the repurposed flow step install all of §4.2; a dind-hosted step cannot write the host routing table (that isolation is the feature), so §4.3 amended the plan: the flow step installs the dind-side legs and a provisioning-time sidecar installs the host route. The
lab-route-installersidecar was chosen overinstall.sh/manual for one-command demo bring-up — and the whole routed path (sidecar, host route, and the flow step's dind-netns legs) was later deleted by the §11 B4 bastion cutover: workers cross the sandbox boundary only through the bastion's authenticated listeners.Image/build-cache GC. The
sandbox_image_pruningmaintenance job sweeps dangling images and build cache on the attested sandbox daemon every 6h by default (§6.4). Reclaiming tagged unused images has no automatic answer — the daemon exposes no last-used timestamp to age them by — so it stays operator-driven.Image pre-warm.
dind-initpre-warms the egress helper image (§6.1), the one cold pull no user action explains; further pre-warming is a deployment choice via the samedind-initextension point.SSH host-key trust model. Every SSH connection the platform makes passes
known_hosts=None— device sessions (shell_transport.py, netmiko, scrapli) and the bastion hop alike: inventory is the source of truth and host keys are not pinned. That is a platform-wide trust model rather than a bastion-specific gap; pinning (ahost_key_ref/fingerprint field onaccess_config, landing for devices and bastions together) is tracked inTODO.md.
11. Roadmap: Bastion Transport & SOCKS-Aware Probes
Successor to the §4 routed worker→lab data path. Once shipped, the demo's
lab-route-installersidecar, the host route, and the flow step's dind-netnssysctl/iptableslegs are all deleted: the sandbox boundary is crossed only by the bastion's two authenticated listeners — SSH on:2222and SOCKS5 (username/password, RFC 1929) on:1080— published by the dind service and never to the host. Beyond the demo, jump-host support is a general product capability — production management networks are routinely reachable only through bastions.
11.1 Target architecture
ProxyJump mechanics: the transport opens SSH session #1 to the bastion (bastion credentials), requests a direct-tcpip channel to <device>:22, and runs SSH session #2 to the device inside that channel. The bastion authenticates access to the management network and can audit connections; it never sees device credentials or session content. TCP-based probes ride the bastion's SOCKS5 listener instead (no per-probe SSH session) — that listener requires username/password authentication (RFC 1929, credentials from the same vault secret family as the SSH login), so neither published port is an open proxy: both endpoints authenticate, and neither is ever published to the host.
11.2 Device model
access_config gains an optional jump_host block, resolved host-side like the existing ssh section (secret refs included):
access_config:
ssh: { username: "...", password_ref: "{{ secret(...) }}" }
jump_host:
host: dind # any resolvable host; demo uses the dind publish
port: 2222
username: jump
password_ref: "{{ secret('vault://.../bastion/password') }}"
# or key_ref; host_key_policy mirrors the ssh sectionFor fleets, devices reference a shared definition rather than repeating it — the demo's git inventory provider stamps the block on every lab device it yields. Handlers are unaffected: credential/config resolution already happens host-side in HandlerServices.connect() / open_shell() (apps/worker/step_handlers/services.py), so the step SDK ABI does not change.
11.3 Transport support (hegemony-step-plugins)
Transports load from the hegemony.device_transports entry-point group (packages/core/transports/registry.py:26); each wheel adds the hop:
- asyncssh (
transport_asyncssh/.../transport.py:137): open the bastion connection first and pass it astunnel=toasyncssh.connect()— first-class library support. Coversshell.executeand the asyncssh-backed netcli path. - netmiko: establish the paramiko jump channel and hand it over via
sock=. - scrapli: per-transport equivalent (system transport via generated OpenSSH
ProxyJumpconfig; paramiko/asyncssh transports via the channel object).
11.4 SOCKS-aware probes
ICMP cannot traverse an SSH tunnel; TCP can. The bastion therefore also runs a SOCKS5 listener (e.g. microsocks) published as a second port, and the probe layer learns to use it:
BaseProbe.execute(address, options)implementations inhegemony-probe-netgain an optionalsocks_proxyoption (socks5://user:pass-ref@dind:1080, credentials as secret refs resolved host-side — the listener always authenticates):tcp_connectdials through the proxy,http_healthuses a SOCKS-capable client connector,dns_resolveis unaffected (resolution is local), andicmp_pingrejects the option explicitly — a documented limitation, withtcp_connect:22as the recommended liveness substitute for tunneled targets.- The one-shot
probe.*handlers and the background monitor pass the option through their existing config paths (HandlerServices.run_probe, monitor target config) — plumbing, not new machinery.
11.5 Phases
- B1 — asyncssh + schema (platform + step-plugins):
jump_hostschema/validation onaccess_config(_resolve_jump_host_refs, fail-on-malformed so a declared bastion is never silently bypassed), secret-ref resolution inservices.connect()and the shell transport,JumpHostSpecon the SDK spec, asyncsshtunnel=support in both the transport wheel andapps/worker/shell_transport.py, device-form editor block. Unit tests: config round-trip, tunnel establishment against a fake, credential isolation (bastion creds never sent to the device and vice versa), bastion teardown on device-connect failure. - B2 — netmiko/scrapli parity: netmiko rides a paramiko
direct-tcpipchannel intosock=; scrapli rides an asyncssh bastion session into its asyncssh transport plugin viatransport_options["asyncssh"]["tunnel"]. Matrix tests cover password and private-key bastion auth, credential isolation on serialized kwargs, and hop teardown on device-connect failure. - B3 — SOCKS probes:
socks_proxyoption (socks5[h]://[user:pass@]host:port) on the probe layer, backed by a dependency-free RFC 1928/1929 client inhegemony-probe-net—tcp_connectdials through the authenticated proxy,http_healthtunnels its GET,dns_resolveis untouched,icmp_pingrejects the option withtcp_connectas the documented substitute.probe.connectivitydeclares and forwards the option; the background monitor already passes free-form check options through. - B4 — demo cutover (
hegemony-demo-data+ demo overlay +hegemony-inventory-plugins): the lab flow deploysmeridian-bastion(same lifecycle pattern as Gitea:--restart unless-stopped, inner publishes:2222/:1080, sshd and microsocks both authenticated with vault-seeded credentials); the git inventory provider stampsjump_hoston every lab device (the provider wheels now merge that scope fromdefault_access_config); the ops flows' TCP reachability probes ride the SOCKS listener. Thelab-route-installersidecar,HEGEMONY_LAB_ROUTE_SUBNET, and the route-opening step'ssysctl/iptableslegs and their teardown are deleted; §4.4's MTU guidance now applies only to the dind service's own uplink.
Non-goals: latency-sensitive polling through the hop (measure in B4), UDP probes, and bastion HA — one bastion per lab is the demo posture.
12. Roadmap: Flow & Step Egress Policies
Outbound firewall rules for
container.runworkloads, declared per-flow and per-step, edited in the UI, enforced inside the sandbox daemon. The dind relocation is what makes enforcement tractable: rules live in the dind netns, never on the host.
12.1 Model & schema
Two declaration sites, both part of the flow definition (so they ride flow Git-sync, Config Exchange, and the audit trail like everything else):
- Flow level — a top-level
network_policyblock indefinition_json, exactly therun_limitsprecedent (apps/api/schemas.py:2267, stored top-level per:2340):
definition:
run_limits: { max_concurrent_runs: 1 }
network_policy:
egress:
mode: allow-listed # open (default) | deny-all | allow-listed
allow:
- "172.20.30.0/24" # CIDR
- "dind:3000" # stack-service shorthand, resolved at enforce time
- "0.0.0.0/0:443/tcp" # CIDR:port/proto
deny:
- "169.254.169.254/32" # always-deny wins over allow- Step level — flat
egress_mode/egress_allow/egress_denyfields onRunContainerConfig(the pydantic model in thehegemony-steps-containerwheel; flat rather than nested so the schema-driven step editor renders them with the existingx_widget: commandslist widget, no new widget type needed), alongside the flow-settings panel for the flow-level block. The rule grammar is mirrored inpackages/core/egress.py— the platform's own authority for parsing and narrowing comparisons — so the admission gate never imports a plugin wheel.
Resolution semantics (the load-bearing decision): effective policy = flow ∩ step — a step may only narrow what the flow grants; denies are unioned. Widening past the flow policy is rejected at save time by a sibling of the host-access gate, _ensure_run_container_egress_authorized (apps/api/routers/flows/lifecycle.py), which also validates rule grammar; widening requires the admin role. Every host-access field — the full _RUN_CONTAINER_HOST_ACCESS_FIELDS set: network_mode, attach_docker_socket, privileged, pid_mode, extra_mounts (plus network_disabled) — makes a step sandbox-admin or hands it firewall-bypass capability (§6.3; a privileged container can manipulate its own network stack, arbitrary mounts can reach the dind control surface), so the same gate rejects at save time (a 400 for everyone, admins included) combining any of them with an egress policy rather than pretending to enforce it; the worker's dispatch re-checks the same invariant and fails closed as a backstop.
Which nodes the gate covers. Admission asks the API's handler registry whether a step launches a container (container.run plus every handler whose wheel declares container_backed). That answer is local: for a handler whose wheel the API lacks or failed to import it would be "no", while a worker that loads the wheel runs a container with whatever was stored. So the gate does not trust the handler id alone. The egress and host-access fields are admitted by name on every node, whatever its handler; the same fields in node.config (which the engine merges under params for every handler but container.run) are a 400 on save and dropped on import; and create, update, duplicate, draft saves and restores of an older version refuse a step whose handler the API does not know, or a container option its handler does not declare (the tool handlers take no host-access options, so such a key would do nothing but fail the step at dispatch). Both are a 400; imports fail closed with the same reason. Flows already stored stay readable.
12.2 Enforcement (dind mode)
- The worker resolves the effective policy per step, canonicalizes it (sorted rules, each reduced to one spelling per meaning — host bits masked, singleton port ranges collapsed, IPv6 literals compressed — so
10.0.0.1/24and10.0.0.0/24cannot provision two networks holding identical rules), and derives its identity as a SHA-256 digest; names embed a 16-hex prefix, but the prefix is never the security boundary: the network is labeled with the full digest and the canonical policy JSON, and before reuse the worker compares the stored canonical policy byte-for-byte — a mismatch (prefix collision or manual tampering) fails the step rather than merging isolation domains. Per distinct policy it ensures, on the dind daemon:a dedicated bridge network
hegemony-egress-<digest16>created withcom.docker.network.bridge.enable_icc=false— but ICC-off is not relied on alone (a broad allow ACCEPT is terminal in FORWARD and would shadow Docker's own intra-bridge drop), so the rule block below also owns an explicit container-to-container drop; concurrent runs on the same policy network therefore cannot reach each other, andDOCKER-USERrules in the dind netns, in this top-down order:-d <subnet> -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPTfirst (return traffic enters the subnet as destination, and the-dscope means this can never accept a NEW outbound flow a conntrack helper marked RELATED); then-s <subnet> -d <subnet> -j DROP(the container-to-container isolation drop); then one drop per deny entry; then, forallow-listed, one accept per allow entry and finally the-s <subnet>default drop. These live in a per-policy chain (HEG-<rules-digest12>), reached from twoDOCKER-USERjumps —-s <subnet>for outbound and-d <subnet>for replies — each carrying-m comment --comment hegemony-egress-<digest16>so ownership, idempotent insertion, and cleanup all key off the jump (the retire loop readsiptables -S, discards the-A DOCKER-USERprefix, and re-issues-Don the bare spec). The chain is named for the rules it holds, so an unchanged policy re-asserts to a no-op, and a changed one is built in a fresh chain that a single atomic jump replacement switches to — a packet is therefore always evaluated against one complete rule set, never a partial one and never none. Superseded chains are reaped once nothing jumps to them.the same chain reached from
INPUTas well, plus an unconditional drop on the daemon's own API port.DOCKER-USERhangs offFORWARD, so it only ever sees traffic routed through the dind host. A packet addressed to the host itself — the policy bridge's own gateway address, or any address the dind container holds — is delivered locally and traversesINPUT, whichDOCKER-USERis never called from. That is not a corner case: dockerd's API listens on those addresses (tcp://0.0.0.0:2375in the overlay), so without anINPUTguard an allow-listed container dials the Docker API straight past its own default drop and creates an unpoliced — or privileged — sibling container. PointingINPUTat the policy chain (-s <subnet>and, when known,-i <bridge>;-d <subnet>is meaningless for host-addressed traffic) makes one rule set govern both paths. The API port additionally gets its own drop above those jumps, because anopenpolicy's chain has no default drop and reaching the daemon is privilege escalation rather than egress — no policy of any mode may permit it. The port comes fromHEGEMONY_CONTAINER_DOCKER_HOST; aunix://endpoint needs no guard, being unreachable from a policy network to begin with.These guards are scoped to policy subnets and bridges, and that scoping is deliberate rather than incidental. A step with no egress policy runs on ordinary Docker networking and keeps its reachability of the sandbox daemon — that is the trust boundary the overlay documents ("any
attach_docker_socketstep is a sandbox admin by design"). Declaring an egress policy is precisely the act of placing a step outside that boundary, so it is exactly the policied steps whose daemon access must go.
- Rules are programmed through a short-lived helper container the worker runs on the dind daemon (
--network host --cap-add NET_ADMIN). The helper uses the same pinneddocker:28-dindimage as the daemon itself, so its iptables binary and backend (legacy vs nftables) cannot diverge from the tables dockerd writes. That image is a host-side image for the dind service but a pull for the sandbox daemon, sodind-initpre-warms it (§6.1) and the worker announces it as run events if the pre-warm did not take. The helper additionally sanity-checks that theDOCKER-USERchain exists before inserting (its absence means the daemon is not managing iptables — fail the step, do not create the chain ad hoc). Inside the sandbox this is a sandbox-level privilege, not a host one. The result is cached per digest; because dind-netns rules evaporate on a dind restart, the worker re-asserts before each policied step launch — which costs one-Nattempt plus a-Cprobe per jump when nothing changed. - The handler attaches the step container with
--network hegemony-egress-<digest16>:ContainerRuntimegains a per-stepdocker_network, populated byWorkerHandlerServices.container_runtime()(apps/worker/step_handlers/services.py:274) — the same settings-channel pattern asdocker_host, one more wheel release.mode: deny-allshort-circuits to the existing--network=none(network_disabled), no rules needed. - DNS is an explicit, documented policy exception in v1, not an accident: Docker's embedded per-network resolver (127.0.0.11) is proxied by the dind dockerd itself and does not traverse the FORWARD path, so name resolution works under default-drop while every resolved connection is filtered. The consequence is stated plainly in the policy docs and the UI helptext: a restricted container can still emit DNS queries through the daemon's resolver (a low-bandwidth exfiltration channel). Tests cover both halves — resolution succeeds, the connection to a non-allowed resolved address is dropped. Closing the exception (an enforced in-sandbox resolver, or
dns: nonefor restricted policies) and hostname-rules are deferred to the egress-proxy phase that can reuse the §11 SOCKS infrastructure; CIDR/port/proto only in v1. - Enforcement runs only against a verified sandbox daemon, and fails closed everywhere else.
dind-initstamps the daemon at provision time with a marker volume whose label carries a deployment-scoped attestation token (hegemony.sandbox.token=<value>, sourced from a compose-provided secret env shared with the worker) — not a barehegemony.sandbox=trueboolean, which anyone able to create a volume on any daemon could spoof. Before programming any rule the worker compares the marker token against its own copy. A step whose effective policy is notopenfails with a clear error whenHEGEMONY_CONTAINER_DOCKER_HOSTis unset (host-socket mode) or when the configured endpoint lacks the marker or carries a mismatched token — so pointing the setting at the host's own socket or an arbitrary TCP daemon can never cause the platform to mutate non-sandbox firewalling, and unenforced policies never silently pass. (Scope honestly stated: this is misconfiguration armor, not defense against an attacker who already controls the worker's env or the target daemon — such an attacker owns the deployment regardless.) Admission still validates the schema everywhere. Tested for all four endpoint shapes: unset, marked-and-matching, marked-but-mismatched, unmarked.
12.3 Phases
E1 — model + admission + UI:
network_policyflow block +FlowNetworkPolicyschema,egress_*fields onRunContainerConfig, the sharedpackages/core/egress.pygrammar, narrowing-only validation inlifecycle.py(_ensure_run_container_egress_authorized), flow-settings panel and schema-driven step-editor widgets,task api:regeneratefor the OpenAPI/UI types. Ships validation-only (no runtime change), independently releasable.E2 — dind enforcement (
apps/worker/egress_enforcement.pywith thepackages/core/egress.pyresolution/digest helpers; the step dispatch inflow_activities.pyfails container steps closed and hands the prepared network to the handler viaContainerRuntime.docker_network): policy-network provisioning (ICC off, digest-labeled, canonical-policy verification before reuse) + rule programming + re-assert logic in the worker; the tokenized sandbox marker indind-init;docker_networkplumbing through the SDK/handler (wheel bump); a janitor sweep for orphanedhegemony-egress-*networks; fail-closed behavior for unset, unmarked, and token-mismatched endpoints. Unit tests: resolution/intersection semantics, digest determinism and collision-rejection, generated rule text (ESTABLISHED/RELATED accept ordered before the default drop), argv (--networkpresent, argv otherwise invariant), fail-closed paths, and the negative space — deny-all short-circuits to--network=nonewith no rules programmed, a policied step against an unmarked or mismatched daemon fails whileopensteps on the same daemon still run, a canonical-policy mismatch on a reused network name fails rather than attaches, the janitor removes only digest-labeled orphans, and a missingDOCKER-USERchain aborts instead of creating it. Live smoke: an allow-listed step reaches only its allow list (and its replies come back); two containers on one policy network cannot reach each other; DNS resolves while the resolved connection is dropped; a second policy gets a second network; dind restart re-asserts.E3 — observability + docs: after every policied step the worker reads per-rule packet counters off the sandbox and surfaces one run event — "N packet(s) dropped, M accepted" with per-rule detail — so a blocked connection is a visible line in the step's event stream instead of a silent timeout.
The enforcement rules are scoped to the policy subnet, so their counters belong to every step under the policy at once. Serialized steps hid that behind a before/after difference; concurrent ones cannot, because their windows overlap — a parallel loop's four iterations each reported all four iterations' traffic as their own. So each step's container launches with a hardware address derived from its own step run id (
container_mac,--mac-address), and a counting-only mirror of the policy rules scoped to that MAC (HEGC-<hex>, see_accounting_specs) is installed above the policy jump. Every rule in the mirror targetsRETURN, so a wrong mirror costs accuracy and never containment, and the step is still enforced by the shared chain below it. The chain is flushed as it is installed, so the read needs no baseline — and the read retires the chain in the same helper invocation, leaving a policied step no more sandbox round-trips than before. A step with no run id falls back to the shared chain and to a report that says the numbers are policy-wide.The identity is a MAC because an address had to be chosen. Docker rejects
--ipon a network whose subnet its own IPAM assigned ("user specified IP address is supported only when connecting to networks with user configured subnets"), and only the IPv6 ULA was named — so the first shipped form of this asked for an address on every policy network and failed every policied container step. Naming a v4 range up front is not an option either: nothing here can see the host's routing table, and a range the operator already routes would blackhole traffic silently. Adopting docker's own choice (inspect, delete, re-create naming the same range) worked, but left a network needlessly re-created and an address pool needing an allocator — and the allocator's record of what was handed out was per process, whiledeploy.replicas: 2against a singledindis the dev and demo topology (§13). Two replicas drawing from one subnet could pick the same address, and the loser'sdocker startfailed with "Address already in use" — a failed step, not a degraded report.A MAC removes all of it. Docker accepts a fixed MAC on a network whose subnet it assigned itself, so nothing is pinned and no network is re-created; deriving it from the step run id means nothing is allocated, nothing is reserved, and two replicas computing it independently cannot collide. It is also the harder of the two to forge from inside the container: rewriting an interface's MAC needs
CAP_NET_ADMIN, which docker withholds by default, while forging a source address needs onlyCAP_NET_RAW, which it grants. A step deliberately givenNET_ADMINcan claim another step's counters, but such a step can already do worse.What a MAC cannot express is a destination:
-m macreads the source hardware address of a frame and has no counterpart. So the mirror counts what a step sent, replies stay with the shared chain, and the per-step report labels its scope "sent" rather than implying both directions.Residual caveats: a concurrent launch that changes the resolved policy rules re-inserts its jump above everything, so containers already running are then decided by the policy chain before their mirror sees them (an undercount, never a difference in enforcement).
Accounting chains are swept on every launch, which they have to be: a MAC derived from a step run id is never reused, so nothing recycles a chain and the happy-path retirement covers only steps that reported. A chain survives while its container is attached to the policy network (the daemon's answer, not a worker's memory, so replicas do not disturb each other) or while it is younger than a grace period, which covers the gap between programming the policy and the container attaching — a gap that contains the step image's pull. Deleting one that is still wanted costs attribution and nothing else, which is why this sweep runs per launch while the policy chains' reap stays the startup janitor's.
The report is prose and data.
format_egress_reportreturns both, so the run event carriesattrsnaming every rule — its role, family, destination, port qualifier and packet count — and diagnosing a blocked connection does not mean parsing a sentence. Rules that matched nothing are in the data even though the prose omits them: "the allow rule for this destination matched zero packets" is frequently the whole diagnosis.format_worker_egress_reportdoes the same for the gate's refusals, which it was already deduplicating and discarding, so both halves of one policy are reported in comparable shapes.Three places, deliberately: the run event for watching it happen,
StepOutput.metrics["egress"]for the totals (scope, dropped, accepted, refused — nothing more, becausemetricsis written verbatim into both the "Step completed" event andStepRun.result_json), and a run artifact (ArtifactKind.EGRESS_REPORT, migration 038) holding the full per-rule detail for the investigation that happens afterwards. The artifact is redacted with the step's resolved secrets on the same terms as evidence — a refusal names the destination it refused, and a destination can come from a secret.Per-destination drop detail: each drop-role rule in a step's accounting mirror carries a rate-limited NFLOG sampler ahead of it (same match, non-terminal), tagged
<digest8>:<mac>via--nflog-prefix. Theegress-nflogsidecar — ulogd2 withnetwork_mode: "service:dind", because nfnetlink_log groups are netns-scoped — writes each sampled packet as one JSON line on a shared volume, and the worker reads the tail back after the step, folding "denied destinations (sampled)" into the report's prose, attrs and artifact. The correlation key travels in the prefix rather than the packet's MAC header deliberately: ulogd 2.0.8's HWHDR interpreter has an off-by-one heap overflow that aborts on the first packet (reproduced under fortify), and the prefix needs no interpreter at all. Sampling means these numbers legitimately disagree with the exact rule counters — every surface labels them sampled. Losing the sidecar degrades, never breaks: NFLOG with no listener still falls through to the counting RETURN, andHEGEMONY_CONTAINER_EGRESS_NFLOG_GROUP=0rolls the rules back to byte-identical pre-feature form.Imports are admitted too, not only API writes: the sync adapter and git pull carry no principal, so
clamp_container_egress_for_import(apps/api/services/egress_admission.py) treats repository content as non-admin-authored — widened step params are narrowed at the same choke point every import funnels through, before diffing so sync plans do not churn, and every dropped rule is reported (pull history, import warnings, logs). Grammar errors and host-access conflicts fail the import closed, as the API gate fails them for every principal.Future: org/global default policies layered above the flow level (same intersection semantics, settings- or org-scoped storage), and domain-level rules via an egress proxy (see §14 for the eBPF alternative).
12.4 Enforcement limits (v1, stated honestly)
- Dual-stack. Rules are programmed with
iptablesfor IPv4 andip6tablesfor IPv6 (bracketed grammar for v6 destinations,[2001:db8::/64]:443/tcp— a bare colon is ambiguous inside a v6 address). Policy networks are created dual-stack (--ipv6plus a deterministic ULA /64) whenever the daemon accepts it — also for v4-only policies, so attaching to a policy network never strips a container's IPv6. Each family is then governed by its own table:openmode leaves a family with no denies fully open ("allow all except 1.1.1.1" keeps v6 reachable);allow-listeddefault-drops every family, so a family the policy never names is closed rather than an escape hatch; ESTABLISHED/isolation/default-drop are scoped per family's subnet. On a daemon without ip6tables the create falls back to v4-only — fine for a v4-only policy, while a policy naming v6 destinations fails closed (half-enforcing is not enforcing). The dind daemon runs with--ipv6 --ip6tablesso both chains exist. - Embedded-DNS side channel. Containers resolve via Docker's embedded resolver (127.0.0.11); dockerd forwards those queries from its own netns through OUTPUT, which
DOCKER-USER(FORWARD hook) never sees. Anallow-listedcontainer whose traffic is fully dropped can still resolve arbitrary names — a low-bandwidth exfiltration channel.deny-allis immune (--network nonehas no resolver). Closing it means pinning policied containers to a controlled resolver (--dns) plus allowing only port 53 to it, or OUTPUT-chain rules in the dind netns. - Published-port targets match post-DNAT.
DOCKER-USERsits in FORWARD, after PREROUTING DNAT. A rule allowing a service published by another in-sandbox container (e.g. Gitea's-p 3000:3000, reached asdind:3000) resolves to the dind service's compose address — but by FORWARD time the destination is already rewritten to the target's inner address. Allow rules for in-sandbox published services must therefore cover the target's inner network (for the demo lab:LAB_MGMT_SUBNET), not thedindcompose address;service:portshorthands are for worker-side or external targets. Note the flip side: such an allow ACCEPT is terminal in FORWARD and intentionally punches through Docker's inter-bridge isolation for exactly that destination. - Point-in-time service resolution.
service:portshorthands resolve once, at enforcement, in the worker's netns. A target that changes address mid-step keeps its old /32 until the next policied launch. - Open mode shares the default bridge. Only policied steps get a dedicated ICC-off network; concurrent
opensteps land on the dind default bridge and can reach each other. - The trust boundary is the workload. A step granted
attach_docker_socketis root over the sandbox daemon; admission refuses to combine it with an egress policy, but it can in principle undermine other containers' networking. Policies defend against code running inside a container, not against authors handed the daemon. - The daemon's own API is closed by INPUT rules, not by
DOCKER-USER. With the default TLS-off TCP listener (§6.3) the sandbox daemon answers unauthenticated on2375— including on each policy bridge's gateway address, which is a local address of the dind netns. Traffic to a local address takes the INPUT path, and the policy rules above hang offDOCKER-USER, a FORWARD hook, so no-s <subnet>deny or default drop there can match it; a policied container reaching2375could create aNetworkMode=hostsibling and egress freely, or flush the policy block outright. Enforcement therefore also installs INPUT jumps and an unconditional API-port DROP for every policy subnet and bridge (§12.2), which is what makes the policy hold. The residual exposure is unchanged for unpolicied steps: everything on the compose network is inside the sandbox's trust boundary by design, and closing that would mean the unix-socket transport (§6.3) or TLS with client certs. - Distinct-policy ceiling. One bridge network per distinct effective policy; Docker's default address pools allow roughly 30. The janitor sweeps unattached policy networks at worker startup only, retiring each removed network's
DOCKER-USERjumps and policy chains with it — Docker recycles subnets, so firewall state left behind would be inherited by whichever network next lands on the freed range. Raising the ceiling =default-address-poolsin the dind daemon config plus a periodic sweep.
Interaction with §11: bastion traffic originates from the worker process, not from step containers, so the kernel rules above never see netcli/probe paths — those are governed by the worker-side gate (§12.5) instead. A policied container step that needs the demo Gitea allows the lab management subnet on port 3000 (see the post-DNAT note above for why not dind:3000).
12.5 Worker-step policy gate
Flow-level egress rules also apply to worker-side steps (netcli, shell, probes, monitors) — with a deliberately different mechanism and a deliberately different trust claim.
Why not kernel rules. Worker handlers run inside the worker process: one netns shared by concurrent flows with different policies, no NET_ADMIN, and the worker's own infrastructure traffic (Temporal, API, the secret store, DNS) that no flow policy may constrain. Per-step kernel rules in that netns are unsound by construction.
What is provided instead: policy-checked connections. The step dispatch binds the flow's effective policy in a contextvar (apps/worker/egress_gate.py, symmetric with the container path's docker-network channel), and the platform-owned connection brokers evaluate every destination before dialing:
| Broker | Covers | Evaluated |
|---|---|---|
HandlerServices.connect | netcli/cisco device CLI | device mgmt_host:mgmt_port and the bastion hop, before any secret resolution |
open_shell_transport | shell steps | same two hops |
HandlerServices.run_probe | probe.connectivity one-shots | target address + check port/proto, and the socks_proxy endpoint when one is set — the proxy is the connection the worker actually opens and authenticates to, so it is governed like the SSH bastion hop; proxying never launders a denied target |
HandlerServices.start_monitor | monitor.* background monitors | every resolved target at start (targets are fixed for the monitor's lifetime; the loop's asyncio task carries only a frozen copy of the step's policy context — created inside the step and outliving it — so a per-tick re-check would look live while being unable to observe any change). Kernel parity: a denied target does not refuse the monitor — it is carried as blocked, never dialed, and reports a failed policy_denied sample naming the rule every tick, exactly as a container pinging a denied address keeps running while those pings fail. The target set a join gate waits on is never thinned; a blocked target simply never comes up. |
Observability, symmetric with §12.2. Every refusal the gate issues during a step is recorded (contextvar-scoped, like the policy binding) and surfaced after the handler as a run event — Egress report: the flow egress policy refused N connections — … — the worker-side analog of the container report's packet counters. A blocked monitor target additionally reports a failed policy_denied sample naming the rule every tick, and the monitor's end emits the totals next to its completion event — Egress report [monitor 'X']: N of M probe(s) blocked by the flow egress policy — <target>: Nx blocked (<rule>).
Evaluation reuses the exact rule semantics the kernel side emits (packages.core.egress.rule_matches_address mirrors the iptables specs: port-less rules match any port/proto, ported rules match only tcp/udp dials in range — so deny 1.1.1.1 blocks a ping while allow x:22/tcp does not admit one). service:port shorthands resolve against the worker's resolver, both families. Hostname targets resolve here too, and every resolved address must pass — the transport may dial any of them. Unresolvable names under a restrictive policy fail closed. dns_resolve checks are exempt (they ask the system resolver about a name; nothing dials the target).
The trust claim, stated honestly. This is enforcement against configuration — where a step is pointed — not against malicious code. Worker handlers are trusted platform code (the §12.4 boundary: "policies defend against code running inside a container"); a userspace gate is the right bar for them, and untrusted workloads belong in container.run where the kernel rules apply. probe.http and probe.wait_reachable dial inside their wheel rather than through run_probe, which is where the gate lives — so they used to bypass it entirely, and a restricted policy did not restrict them at all. They now call HandlerServices.enforce_egress(host, port) immediately before dialling: the same denial evaluation run_probe performs, at the seam a handler that owns its own socket can reach. wait_reachable vets every target once before its poll loop rather than per poll, so a denial fails the step with the policy's reason instead of after max_wait_seconds of waiting on a host the flow may not reach. Any future handler that opens its own connection has the same obligation. Notification dispatch and internal API traffic remain deliberately exempt (admin-configured infrastructure, not flow-authored targets); monitor target resolution happens at start, so a DNS name that moves mid-monitor keeps its start-time verdict.
13. Multi-Worker & Remote-Worker Topology
Decision: one dind per host — every worker on a host shares that host's single sandbox daemon; a host without a dind runs no container steps, and there is never more than one dind per host. This is neither one-per-worker (which would forfeit
deploy.replicas) nor one central dind for the whole deployment (a dind is host-local and cannot be shared across hosts). The shipped stacks already embody it: dev, prod, and the demo rundeploy.replicas: 2workers against the singledindthe overlay adds.
13.1 Why per-host
- The host is the natural boundary. A host's worker replicas already share its kernel, its host-local
shared-workspacesvolume, and — in the demo — the lab the flow deploys. Binding the sandbox to the host rather than the worker process is what keepsdeploy.replicas(no one-service-per-worker) while letting every replica resolve the samedindby compose DNS with the shared workspace already coherent. - Not one-per-worker. Workers on a host sit inside one trust boundary — any
attach_docker_socketstep is a sandbox admin whichever replica runs it — so a private dind per replica buys no isolation. It would only duplicate the image store and the digest-named policy networks on the same host, and force one compose service per worker. - Not one central dind. A dind is host-local by construction: privileged, sharing the host kernel, holding a host-local data volume and image store, and hosting the bastion the lab flow publishes in its own netns. It cannot be mounted or reached across hosts, and a single fleet-wide dind would be both a cross-host network dependency and a single point of failure for every container step everywhere.
- Concurrency across a host's workers is expected — and handled. N replicas program one daemon at once, which is why the enforcement is race-safe: policy chains are content-named and published by an atomic rename (
iptables -E, §12.2), and the startup janitor (apps/worker/container_cleanup.py) refuses running, recently-created (grace window), and foreign-stack containers, so one replica never corrupts another's rules or reaps its in-flight step. Per-host sharing is the reason that machinery is load-bearing rather than incidental. - Blast radius is the host. A runaway step that wedges or fills a dind takes down container-step capacity for that host's workers only; other hosts keep executing. Coarser than a per-worker daemon would give — a host's replicas share fate for container steps — but bounded, never fleet-wide.
- Attestation already fits N hosts. The sandbox token is deployment-scoped, not daemon-scoped: every host's dind is provisioned (via its own
dind-init) with the sameHEGEMONY_SANDBOX_TOKEN, and each worker attests its own host's daemon. No schema change for more hosts.
13.2 Remote hosts
A remote worker host deploys the same shape: its own dind on a host-private network, HEGEMONY_CONTAINER_DOCKER_HOST pointing at that host's daemon, which must never be reachable from anything but that host (TLS-off TCP is the current transport — §6.3; keeping the daemon host-private is what makes that acceptable, remotely as locally). The worker's only cross-host dependencies stay what they are today: Temporal, the API, and the secret store. Each host's dind needs registry egress (or a per-host mirror) to pull step images.
Two consequences to design around:
- Shared workspaces and dind-local state are per host.
/data/shared, the image store, any/tmpa flow writes inside dind, and the lab bastion all live on one host's daemon. A run that shares files between steps, builds an image in one step and consumes it in a later one, or reaches the lab, must run all its steps on one host. Run affinity already provides this:execution_affinity: shared/explicitpins the whole run to one worker — hence one host, one dind, one volume — by routing its steps to that worker's per-host Temporal queuehegemony-host-<id>(apps/worker/run.py,apps/worker/flow_workflow.py). On a single-host stack every worker shares the one dind, so pinning is belt-and-suspenders there; it becomes necessary only once hosts differ. Long term, artifact-store handoff (the S3 object store, already in the stack) removes the bind-mount coupling for pure file passing. - Per-host bastions. §11 listeners publish in each host's dind netns, so lab access is inherently per host — a run whose steps need host A's lab must route to host A anyway. Affinity is therefore not just a workspace concern; it is the correct scheduling model for multi-host labs.
Rollout sketch: the worker gains an optional host/site label; the API maps flows (or inventory scopes) to queues; the demo is a single host — one dind, two worker replicas — and is unaffected. One caveat for a future multi-host rollout: flows that lean on dind-local state today without declaring affinity (the demo Lab flow builds images on one step and deploys them on another, correct only because a single host means a single image store) would then need execution_affinity: shared, or a registry both hosts pull from.
14. eBPF Egress Enforcement: Feasibility
Assessment: possible, not warranted for v1. iptables in the dind netns stays. Revisit when hostname-level rules become a requirement. (Per-step attribution no longer motivates it: E3's per-container accounting chains give it within iptables.)
Possible. The dind service is privileged and shares the host kernel, so cgroup-attached BPF programs (connect4/sendmsg4 hooks, or tc/TCX on the policy-bridge interfaces) can be loaded from inside the sandbox. Per-policy allow/deny maps keyed the same way as today's digests would replace the DOCKER-USER blocks one-for-one, with per-map counters giving the same observability without a helper container.
Why not now:
- Parity, not capability. For pure L3/L4 CIDR/port/proto rules, eBPF adds no expressiveness over the shipped iptables blocks — it would re-implement a working, tested enforcement path.
- The features that would justify it are bigger than the loader. Hostname/SNI-aware policies need DNS-answer capture correlated with connect destinations (the hard part of Cilium-class engines), or a transparent egress proxy — which delivers the same user-visible feature with far less machinery (§12 Future).
- Operational surface. Kernel-version sensitivity (the sandbox shares whatever kernel the host runs), CO-RE toolchain and vmlinux BTF availability inside
docker:28-dind, program lifecycle across dind restarts, and debugging opacity for operators — all new failure modes replacing well-understood iptables semantics. - The kernel is shared. BPF loaded "inside the sandbox" still runs in the host kernel. That widens the blast radius of an enforcement bug from "rules in a throwaway netns" to "programs in the shared kernel" — the opposite of the sandbox's isolation direction.
If adopted later: keep the policy model, digests, and admission unchanged; swap build_iptables_script/helper-container for a pinned loader image (BTF-checked at attest time) that populates per-digest maps; keep DOCKER-USER programming as the fallback path behind the same fail-closed gate, selected per daemon by capability probe at attestation. The contextvar/docker_network handoff is enforcement- agnostic and would not change.