Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 23 additions & 13 deletions .agents/skills/trust-boundary/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@ description: >
Checklist for code that decides which devices a container may open, or what the
daemon believes about its own configuration. Apply before editing internal/config/**,
internal/policy/**, the label handling and rule collection in internal/processor/**,
the SIGHUP reload path in cmd/swarm-device-access/config.go, deployments/** (including
the Dockerfile) or examples/**. Read it before writing the change, not after review
the SIGHUP reload path in cmd/swarm-device-access/config.go, internal/launcher/**,
deployments/** (including the Dockerfile) or examples/**. Read it before writing the change, not after review
finds the hole.
---

Expand Down Expand Up @@ -161,18 +161,28 @@ file are logged and ignored. Covered by `TestReload_*` and `TestMergeSettings_*`
## 6. Least privilege in the image and the deployment

The daemon has to be privileged: it attaches BPF programs to other containers' cgroups and stats
host device nodes. So the rule is not "unprivileged" but "no more than the documented set". The
README's Docker Compose for Swarm snippet and `deployments/docker/docker-compose.yaml` run it with
`--privileged`, `--cgroupns=host`, `--pid=host`, `--userns=host` and exactly three bind mounts:
`/sys:/host/sys`, `/var/run/docker.sock` and `/dev:/dev`. The DBus socket
(`/run/dbus/system_bus_socket:/var/run/dbus/system_bus_socket`) stays **optional and commented
out** — the daemon degrades to a startup warning without it. `deployments/docker/Dockerfile` builds
a static binary onto `dhi.io/static` with `USER 0` and nothing else in the runtime stage.

- [ ] No new capability, namespace flag or host bind mount beyond the documented set, in the
Dockerfile, `deployments/**`, `examples/**` or the README — and if one is truly needed, the
host device nodes. So the rule is not "unprivileged" but "no more than the documented set".

The single source of truth for that set is `daemonSpec` (`internal/launcher/spec.go`): the Swarm
service in the README's Docker Compose for Swarm snippet, `deployments/docker/docker-compose.yaml`
and `examples/**` runs the image as an unprivileged `launch` task with only the Docker socket, and
the launcher creates the daemon with `Privileged`, host cgroup, PID and user namespaces, network
`none`, `AutoRemove`, and exactly three bind `Mounts`: the Docker socket at `/var/run/docker.sock`,
`/sys:/host/sys` and `/dev:/dev`. `Mounts`, never `Binds`: a missing source must fail the create,
not make Docker create an empty directory on the host. The optional extras are opt-in launch flags:
`-dbus` (`/run/dbus/system_bus_socket:/var/run/dbus/system_bus_socket` — the daemon degrades to a
startup warning without it), `-config-dir` (read-only at `/etc/swarm-device-access`) and
`-host-network`. `TestDaemonSpec` pins the whole host config, so a new mount or flag shows up as a
test change. The launcher's own inputs are checked too: `-config-dir` and `-host-docker-socket`
must be absolute, clean paths, and it removes a `swarm-device-access` container only when its
owning launcher is confirmed gone — never on a transient error, and never a foreign container.
`deployments/docker/Dockerfile` builds a static binary onto `dhi.io/static` with `USER 0` and
nothing else in the runtime stage.

- [ ] No new capability, namespace flag or host bind mount beyond the documented set, in
`daemonSpec`, the Dockerfile, `deployments/**`, `examples/**` or the README — and if one is truly needed, the
README and `AGENTS.md` runtime requirements change in the same commit.
- [ ] The DBus mount is never made mandatory.
- [ ] The DBus mount is never made mandatory (`-dbus` stays off by default).
- [ ] A Dockerfile change is checked with `make docker-build`.
- [ ] Nothing secret is baked into a layer: build args and `COPY` sources are not a secret store.

Expand Down
21 changes: 19 additions & 2 deletions .mk/integration-test.mk
Original file line number Diff line number Diff line change
Expand Up @@ -23,12 +23,17 @@
# make go-test-integration SDA_IT_ENFORCE=1 SDA_IT_REQUIRE_RELOAD=1
#
# Every container the suite creates carries the swarm-device-access-it.envid
# label. After a killed run, remove the leftovers of every run with:
# label. The daemon a test launcher creates does not; it is swept once its
# launcher no longer exists. After a killed run, remove the leftovers of every
# run with:
# make sweep-test-leaks

INTEGRATION_TIMEOUT ?= 6m
INTEGRATION_PKG ?= ./test/integration/...
INTEGRATION_ENV_LABEL := swarm-device-access-it.envid
# The labels internal/launcher puts on the daemon containers it creates.
LAUNCHER_ROLE_LABEL := io.github.leinardi.swarm-device-access.role
LAUNCHER_OWNER_LABEL := io.github.leinardi.swarm-device-access.launcher

# go test changes CWD to the package dir, so the binary path must be absolute.
export SDA_TEST_BINARY ?= $(REPO_ROOT)/$(DIST_DIR)/$(BIN_NAME)
Expand Down Expand Up @@ -60,4 +65,16 @@ integration-image: ## Build the daemon image the enforcement test runs, as $(INT

.PHONY: sweep-test-leaks
sweep-test-leaks: ## Remove every container left behind by integration test runs
docker ps -aq --filter "label=$(INTEGRATION_ENV_LABEL)" | xargs -r docker rm -f
@set -eu; \
ids=$$(docker ps -aq --filter "label=$(INTEGRATION_ENV_LABEL)"); \
if [ -n "$$ids" ]; then docker rm -f $$ids; fi; \
daemons=$$(docker ps -a --filter "label=$(LAUNCHER_ROLE_LABEL)=daemon" \
--format '{{.ID}} {{.Label "$(LAUNCHER_OWNER_LABEL)"}}'); \
printf '%s\n' "$$daemons" | while read -r id owner; do \
[ -n "$$id" ] && [ -n "$$owner" ] || continue; \
if err=$$(docker container inspect "$$owner" 2>&1 >/dev/null); then continue; fi; \
case "$$err" in \
*"No such"*) docker rm -f "$$id" ;; \
*) echo "inspect launcher $$owner of daemon $$id: $$err" >&2; exit 1 ;; \
esac; \
done
4 changes: 3 additions & 1 deletion .serena/memories/core.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,11 +11,13 @@ cmd/swarm-device-access/ # main binary (//go:build linux)
main.go # run(): flags, settings, openat2 probe, store, processor, daemon.Run
flags.go # CLI flag definitions
config.go # settings merge/validate, SIGHUP reloader
launch.go # `launch` subcommand: launcher flags, runLaunch
version.go # version/commit/date (filled by ldflags)

internal/
config/ # strict YAML loader; runtime Store/Publisher with generations
policy/ # mode, label parsing, globs: Enabled, Denied, Authorized (cross-platform)
launcher/ # launch mode: daemonSpec (privileged set), self ID, stale cleanup, supervise
daemon/ # event loop and coordinator
daemon.go # Run: subscribe, coordinator, startup pass, reload watcher, listenEvents
events.go # listenEvents/consumeEvents (reconnect backoff), processOne
Expand All @@ -38,7 +40,7 @@ internal/
systemd/ # DBus Reloading signal watcher

test/integration/ # daemon integration tests (build tag integration)
deployments/docker/ # Dockerfile, compose file with the sh wrapper, example config
deployments/docker/ # Dockerfile, compose file running the launcher, example config
docs/ # architecture.md, testing.md, release.md
```

Expand Down
15 changes: 12 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,14 @@ and aggregates per device), then apply it (`applyPinned`: pidfd pin, `/proc/<pid
runtime's `BPF_CGROUP_DEVICE` filter in an owned block and swap it in with `BPF_F_ALLOW_MULTI` (or `BPF_F_REPLACE`). **Parts of this code are
preserved from NVIDIA (Apache 2.0) — touch with care.**

**Launcher (`cmd/swarm-device-access/launch.go`, `internal/launcher/`)**

`run()` dispatches `launch` before `flag.Parse()` to `runLaunch`, which has its own flag set; the arguments after `--` go to the daemon
verbatim. The Swarm service runs the image unprivileged in this mode. `launcher.Run` finds its own container in `/proc/self/mountinfo`,
takes its `sha256:` image ID, removes a stale `swarm-device-access` daemon only when its owning launcher is confirmed gone (or it came
from the old `sh` wrapper), then creates the daemon from `daemonSpec`, attaches, registers `ContainerWait(removed)`, starts it and
supervises it. Stops go by container ID and share one 20 s shutdown budget.

**Glue (`internal/logger/`, `internal/systemd/`, `internal/observability/`)**

- `logger` — slog wrapper with `text`/`json`/`plain` handlers, `-log-time` strips timestamps via a `ReplaceAttr`. `L()` lazy-inits a default INFO text
Expand All @@ -90,8 +98,9 @@ preserved from NVIDIA (Apache 2.0) — touch with care.**

## Runtime requirements

The daemon **must** run with `privileged: true`, `cgroup: host`, `pid: host`, `userns_mode: host`, and bind mounts for `/var/run/docker.sock` and
`/sys → /host/sys`. The `hostRootPath = "/host"` constant in `main.go` is the inside-container view of the host root; cgroup paths (and sysfs, when
The daemon **must** run with `privileged: true`, `cgroup: host`, `pid: host`, `userns_mode: host`, and bind mounts for `/var/run/docker.sock`,
`/sys → /host/sys` and `/dev`. That privileged set lives in one place, `daemonSpec` (`internal/launcher/spec.go`), which the launcher uses to
create the daemon; the Swarm service itself (`launch`) needs only the Docker socket. The `hostRootPath = "/host"` constant in `main.go` is the inside-container view of the host root; cgroup paths (and sysfs, when
mounted there) are joined against it. The DBus socket mount is optional — enables reload handling. Mount as `-v /run/dbus/system_bus_socket:/var/run/dbus/system_bus_socket`; the container-side path must be under `/var/run/` because `dhi.io/static` has no `/var/run → /run` symlink.

## Conventions worth knowing
Expand All @@ -109,7 +118,7 @@ Skills live in `.agents/skills/` (symlinked as `.claude/skills`). Load them befo

- `go-style-guide` — before any `.go` edit.
- `trust-boundary` — before touching `internal/config`, `internal/policy`, label parsing or rule collection in `internal/processor`,
`deployments/docker/Dockerfile` or anything else under `deployments/**`.
`internal/launcher`, `deployments/docker/Dockerfile` or anything else under `deployments/**`.
- `adversarial-review` — for any review request ("review my diff", "is this ready to merge").

## Quality rules not enforced by tooling
Expand Down
137 changes: 70 additions & 67 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,9 +52,11 @@ at `/host/sys`, and `/dev` mounted into the daemon container.
It requires Docker Engine 19.03 or newer (API 1.40): the Docker client it uses
refuses older engines.

Swarm does not allow those runtime options directly on a service. The common
workaround is to deploy a small wrapper service that runs the Docker CLI and
uses the host Docker socket to launch the real privileged daemon container.
Swarm does not allow those runtime options directly on a service. So the image
has a second mode, `launch`: the Swarm service runs the image unprivileged, with
only the Docker socket, and the launcher creates the privileged daemon container
from the very image it runs itself (see
[Docker Compose for Swarm](#docker-compose-for-swarm)).

### Host Requirements

Expand Down Expand Up @@ -163,58 +165,33 @@ is dropped only once its cgroup is verified gone; a periodic sweep checks.
```yaml
services:
swarm-device-access:
image: docker:29
# Swarm rejects privileged/cgroup/pid/userns on services. This wrapper
# launches the actual daemon with `docker run`, where those flags are valid.
entrypoint: ["sh", "-c"]
image: ghcr.io/leinardi/swarm-device-access:1
# Swarm rejects privileged/cgroup/pid/userns on services. The launcher runs
# unprivileged and creates the daemon container from this same image.
command:
- |
# docker run options, before the image. Uncomment a line to add it.
set -- --rm --name=swarm-device-access \
--privileged --cgroupns=host --pid=host --userns=host \
-v /sys:/host/sys \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /dev:/dev
# Reapply device rules after systemctl daemon-reload, which wipes cgroup
# BPF programs (without it the daemon warns and skips reload handling).
# The container-side path must be under /var/run: dhi.io/static has no
# /var/run -> /run symlink.
# set -- "$$@" -v /run/dbus/system_bus_socket:/var/run/dbus/system_bus_socket
# A config file for -config below. Mount the directory, not the file:
# an editor that replaces the file would leave a file mount on the old
# content, and a reload would silently apply it.
# set -- "$$@" -v /etc/swarm-device-access:/etc/swarm-device-access:ro
set -- "$$@" ghcr.io/leinardi/swarm-device-access:latest
# Daemon flags, after the image; before it they are docker run flags.
# Reload the config file with: docker kill -s HUP swarm-device-access
# set -- "$$@" -config /etc/swarm-device-access/config.yaml

# On SIGTERM or SIGINT from Swarm, stop the daemon and exit with its
# status. The trap is set before the daemon starts, and the daemon runs
# in the background: a shell waiting on a foreground child only runs
# its traps after the child exits.
pid=
stop_daemon() {
docker stop -t 10 swarm-device-access >/dev/null 2>&1
wait $$pid
exit $$?
}
trap stop_daemon TERM INT

# Clear a daemon container left behind by a wrapper that was killed.
docker rm -f swarm-device-access >/dev/null 2>&1 || true
docker run "$$@" &
pid=$$!
wait "$$pid"
exit $$?
# Leaves time for docker stop -t 10 before Swarm kills the wrapper.
stop_grace_period: 30s
- launch
# Launcher flags. Reapply device rules after systemctl daemon-reload:
# - -dbus
# Bind a host directory read-only at /etc/swarm-device-access. Mount the
# directory, not the file: an editor that replaces the file would leave a
# file mount on the old content, and a reload would silently apply it.
# - -config-dir=/etc/swarm-device-access
- --
# Daemon flags, passed to the daemon verbatim.
# Reload the config file with: docker kill -s HUP swarm-device-access
# - -config=/etc/swarm-device-access/config.yaml
- -log-level=info
- -log-format=text
volumes:
- /var/run/docker.sock:/var/run/docker.sock
deploy:
mode: global
restart_policy:
condition: any
update_config:
order: stop-first
# The launcher's shutdown takes at most 20s; see below.
stop_grace_period: 30s

# Example consumer service: a Swarm task that bind-mounts a GPU device.
cuda-worker:
Expand All @@ -237,25 +214,51 @@ services:
The daemon compose file is also available at
[`deployments/docker/docker-compose.yaml`](deployments/docker/docker-compose.yaml).

The wrapper is a small `sh` script rather than a bare `docker run`, so that
there is exactly one daemon container per node, always named
`swarm-device-access`:

- It builds the `docker run` arguments with one `set -- "$$@" ...` line each:
options before the image, daemon flags after it. To enable an optional mount
or flag, uncomment its line; it cannot end up on the wrong side of the image.
- It sets a `SIGTERM`/`SIGINT` trap, then removes any `swarm-device-access`
container a killed wrapper left behind, so the replacement task does not
start a second daemon next to an orphan or, normally, hit a name conflict (if
it does, Swarm restarts the task and the next attempt clears it).
- It starts `docker run` in the background and waits on it, because a shell
blocked on a foreground child runs its traps only after the child exits.
- On `SIGTERM` from Swarm the trap runs `docker stop -t 10
swarm-device-access`, waits for the daemon to exit, and exits with its
status. `stop_grace_period: 30s` gives it time to do so.

Compose interpolates `$` in the file, so the script writes `$$` for a literal
`$`.
`swarm-device-access launch [launch flags] -- [daemon flags]` finds its own
container, reads the ID of the image it runs, and creates one container named
`swarm-device-access` from that image ID, with the arguments after `--` as the
daemon's flags. The daemon container gets exactly the documented privileged
set: `--privileged`, the host cgroup, PID and user namespaces, no network, and
bind mounts of the host Docker socket at `/var/run/docker.sock`, `/sys` at
`/host/sys` and `/dev` at `/dev`, plus the optional mounts below. The launcher
streams the daemon's output as its own, so `docker service logs` shows the
daemon, and exits with the daemon's status.

| Launch flag | Default | Description |
| --- | --- | --- |
| `-dbus` | `false` | Bind the host's `/run/dbus/system_bus_socket` at `/var/run/dbus/system_bus_socket` in the daemon, so it reapplies rules after `systemctl daemon-reload`. |
| `-config-dir` | `""` | Host directory bound read-only at `/etc/swarm-device-access`, for the daemon's `-config`. Must be an absolute, clean path. |
| `-host-network` | `false` | Run the daemon in the host network instead of none; only `-metrics-addr` and `-debug-addr` need it. |
| `-host-docker-socket` | `/var/run/docker.sock` | Host path of the Docker socket bound into the daemon. Must be an absolute, clean path. The launcher itself always uses its own `/var/run/docker.sock` mount. |
| `-log-level` | `info` | The launcher's own log level: `debug`, `info`, `warn`, `error`. |
| `-log-format` | `text` | The launcher's own log format: `text`, `json`, `plain`. |
| `-help` | | Print the launch flags and exit |

Running the daemon from the launcher's own image ID, not a tag, means the
service image is the daemon image:

- An image watcher such as gantry, or `docker service update --image`, updates
the service, and the new launcher starts the new daemon. Every node runs the
same digest.
- A Swarm rollback (`docker service rollback`) rolls the daemon back with it.
- `update_config.order: stop-first` stops the old launcher, which stops and
removes its daemon, before the new one starts. The launcher's shutdown (a 10s
stop, the wait for the removal and the drain of the last log lines) shares a
single 20s budget, inside `stop_grace_period: 30s`.

Each daemon container is labelled with the ID of the launcher that owns it.
Before it creates a daemon, a launcher removes a `swarm-device-access`
container only when it is a daemon whose launcher is confirmed gone or
stopped, or a daemon left by the `sh` wrapper of earlier releases. A daemon
whose launcher still runs, a launcher whose state cannot be read, or any other
container with that name makes the launcher exit with an error, and Swarm
retries it.

Releases before 1.0.0 documented a `docker:29` service that ran an `sh` wrapper
around `docker run`. It still works, but it is no longer documented: the
service image never changed, so image watchers and rollbacks did not follow
daemon releases. To migrate, replace the service with the one above; the first
launcher removes the daemon container the wrapper left running.

## ⚙️ Configuration

Expand Down
Loading
Loading