Skip to content

feat(launcher): launch the privileged daemon from the image's own digest - #24

Merged
leinardi merged 6 commits into
masterfrom
feat/launcher
Sep 27, 2026
Merged

leinardi merged 6 commits into
masterfrom
feat/launcher

Conversation

@leinardi

Copy link
Copy Markdown
Owner

Summary

Adds a launch mode to the daemon image, so the Swarm service can run ghcr.io/leinardi/swarm-device-access:1 launch … -- <daemon flags> unprivileged with only the Docker socket. It replaces the docker:29 sh wrapper.

  • Same digest. The launcher finds its own container ID in /proc/self/mountinfo, inspects itself and creates the daemon from its own sha256: image ID. Launcher and daemon therefore always run the same digest, and gantry, rolling updates and docker service rollback act on the daemon as well.

  • One privileged set. daemonSpec (internal/launcher/spec.go) is the single source of truth for the daemon's privileges: privileged, host cgroup, PID and user namespaces, network none (-host-network opts into host), AutoRemove, a 10 s stop timeout, and bind Mounts (no Binds) for the Docker socket, /sys at /host/sys and /dev. -dbus and -config-dir (read-only) are opt-in.

  • Stale-daemon cleanup. Before creating a daemon, the launcher removes an existing swarm-device-access container only in these cases:

    • its owning launcher is confirmed gone or stopped;
    • it was left by an earlier run of this same launcher container;
    • it was started by the old sh wrapper.

    A live owner, a transient inspect error or a foreign container aborts with an error, and Swarm retries the launcher.

  • Supervision.

    • The launcher follows the docker run order: create, attach (the copy goroutine starts before start), ContainerWait(removed), start.
    • The attach and wait streams have their own contexts, detached from the signal. Only their establishment is time-bounded.
    • SIGTERM, a broken wait stream or a lost log stream stop the daemon by ID. The stop, the removal wait and the log drain share one 20 s shutdown budget, which fits inside stop_grace_period: 30s.
  • Tests.

    • Unit tests with a fake Docker API cover:
      • the spec table;
      • mountinfo parsing, including the TrueNAS data-root;
      • the stale-daemon policy;
      • call order;
      • establishment bounds;
      • final-log drain;
      • log-stream and wait-stream loss;
      • cancellation;
      • the shared shutdown budget.
    • TestLauncher_RunsDaemonFromItsOwnImage is an integration test in the enforcement group, so it runs only in CI. It covers:
      • the spec of the daemon container that actually gets created;
      • a graceful stop that exits 0;
      • replacement of a daemon orphaned by a killed launcher.
    • sweep-test-leaks also removes daemons whose launcher no longer exists.
  • Docs.

    • The README, the compose files, the five examples, SECURITY.md, the architecture and testing docs, AGENTS.md and the trust-boundary skill all describe the launcher.
    • The README keeps a note retiring the sh wrapper.

Local verification:

  • go test -race ./..., golangci-lint run (including -tags integration) and make check pass.
  • make go-test-integration passes. It ran dry-run only: the launcher and enforcement tests skip without SDA_IT_ENFORCE=1.
  • make docker-build works. Inside the built image, launch -help exits 0 and launch -config-dir relative fails with a named error. A bare launch identified its own container and then failed only because it had no socket.

Pull request checklist

  • I am targeting the master branch
  • I have rebased this branch on top of the destination branch
  • I have executed make check locally before creating the commit and it has run successfully
  • I have performed a self-review of my own code
  • There are no WIP commits in this PR

Type of changes

  • 🐛 Bug fix
  • ✨ New feature
  • 🔧 Refactoring
  • 📜 Docs
  • 🧰 CI / tooling / infra
  • Other (describe in Summary)

Add a `launch` subcommand so the Swarm service can run the daemon image
itself, unprivileged, with only the Docker socket. The launcher finds its
own container in /proc/self/mountinfo, takes the sha256 image ID it runs
and creates the privileged daemon container from it, so the launcher and
the daemon are always the same digest: image watchers, rolling updates
and rollbacks act on the daemon too.

daemonSpec is the single source of the daemon's privileged set and uses
bind Mounts, not Binds, so a missing host path fails the create. Before
creating, the launcher removes a stale swarm-device-access daemon only
when its owning launcher is confirmed gone or stopped, or when the old sh
wrapper started it; a live owner, a transient inspect error or a foreign
container aborts. It attaches and registers the removal wait before the
start, streams the daemon's output as its own, stops the daemon by ID and
bounds every shutdown by one 20s budget.
Run the daemon image in launcher mode with only the Docker socket and a
dry-run daemon behind it: check the daemon container's privileged set,
image and owner label, a graceful stop that exits 0 and leaves no daemon,
and a killed launcher whose orphaned daemon the next launcher replaces.
It creates a privileged container named swarm-device-access, so it runs
with the enforcement tests.

sweep-test-leaks also removes launcher-created daemons whose launcher
container no longer exists.
The reference compose file, the five examples and the README now run
ghcr.io/leinardi/swarm-device-access:1 as a `launch` service with only
the Docker socket, stop-first updates and a 30s stop grace period; each
example keeps its commented -device-allow lines after `--`. The README
documents the launch flags, why the daemon runs from the launcher's own
image, and retires the docker:29 sh wrapper of earlier releases.

SECURITY.md, AGENTS.md, docs/architecture.md, docs/testing.md and the
trust-boundary skill name daemonSpec as the single source of the
daemon's privileged set.
…r a broken wait

establish waited for the call's goroutine after canceling it. The
client's hijacked attach reads the upgrade response on a raw connection
that no context reaches, so a dockerd that accepts and never answers
kept the launcher blocked forever, with the created container and its
fixed name left behind. It now returns at the deadline and releases a
late result in the background.

A broken wait stream stopped the daemon and returned at once, closing
the attach stream: the last log lines could be cut and the removal was
never confirmed. That path now shares the shutdown budget like the
others: stop, poll inspect until the daemon is gone, drain the output.
The recipe's status came from the last command of each pipeline, so a
failed docker ps, a failed owner inspect or an earlier docker rm still
left make green, with privileged containers behind. List before
removing, run under set -eu, and fail on an inspect error other than
"No such", instead of skipping the daemon.
Polling for the daemon's removal after a broken wait stream dropped
every inspect error but NotFound, so a permission or connection failure
surfaced only as the shutdown budget running out. Return the latest
inspect error with the timeout.
@leinardi
leinardi merged commit ed18618 into master Sep 27, 2026
9 checks passed
@leinardi
leinardi deleted the feat/launcher branch September 27, 2026 13:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant