Repository navigation
Reconcile device grants safely and fix audit findings - #23
Merged
Merged
Conversation
Add DockerCallTimeout next to the backoff constants and apply it to every request/response Docker call: a per-container timeout in processOne, a per-call timeout inside the processor (Processor.CallTimeout, set from DockerCallTimeout) for ContainerInspect and ServiceInspect so every entry point is bounded, every ContainerList (startup and systemd re-apply enumeration) and the startup Info call. Event stream establishment is bounded in the caller: moby's Events blocks until dockerd sends the response headers, so it runs in a goroutine on a cancelable stream context and is raced against a timer after a bounded Ping pre-check. On timeout the request is canceled and the goroutine awaited, and the loop re-enters the backoff; on success the stream context lives until reconnect or shutdown. A failed initial subscription no longer blocks startup: listenEvents resubscribes from the original since.
Policy now comes from the global config and the container's own labels only, so the decision for a container is identical on every node and across daemon restarts. Service-level labels were visible only on managers, unenforceable on workers and silently dropped when the service inspect failed, which turned a lost narrowing label into wider access. Removed: resolveServiceLabels, policy.MergeLabels, policy.KnownLabels, Processor.IsSwarmManager, the Info-based Swarm role detection at startup and the opt-in-granted-via-service-level-label log. ServiceInspect leaves the DockerInspector interface; the processor makes no service or node inspection. Unknown swarm-device-access.* keys are now reported on the container labels, the only label input left. BREAKING CHANGE: labels under deploy.labels are ignored; put swarm-device-access.* labels under the service's top-level labels: so they reach the container. A service whose only opt-in is a deploy-level label is now skipped.
Add one Processor-wide mutex held from the config load through the cgroup mutation, not around the mutation alone: the event loop, the startup enumeration and the systemd re-apply can process the same cgroup concurrently, and a worker that computed rules under an old config must not apply them after a newer config was loaded and applied by another worker. Docker calls made under the lock stay bounded by CallTimeout. A test holds one worker between compute and apply, starts a second worker for the same container, publishes a narrower config, and checks the second neither computes while the first holds the lock nor sees the old config afterwards.
…window AddDeviceRules detached every attached program before attaching the replacements, as NVIDIA's code does because it runs before the container starts. This daemon mutates live containers, so that left an unfiltered window, and a failed attach or daemon death left the container unfiltered for good. Programs are now replaced pairwise: atomically with BPF_F_REPLACE (via cilium's ReplaceProgram anchor) when the kernel supports it, probed on first use and cached, otherwise by attaching the new program before detaching the original. A detach failure rolls that pair back; earlier pairs stay replaced, which under BPF_F_ALLOW_MULTI is narrower, never wider. FindAttachedCgroupDeviceFilters now returns the kernel's total, the inaccessible count and the attach flags, and closes opened handles on every error path. A filter attached without BPF_F_ALLOW_MULTI skips the container (unsupported_attach_mode); any inaccessible filter is a retryable ErrFiltersInaccessible with nothing mutated; no attached filter is ErrFilterMissing instead of the Return()-only fallback the verifier always rejected. Also: RLIMIT_MEMLOCK is removed once at startup with rlimit.RemoveMemlock, runtime.KeepAlive covers the query buffer right after the syscall, logging goes through logger.L, and the kernel calls sit behind an ops seam with order tests. BREAKING CHANGE: containers whose runtime attaches the device filter without BPF_F_ALLOW_MULTI, and hosts where an LSM policy hides a device filter from the daemon, are now refused (skipped or retried without changes) instead of handled best-effort. See Host Requirements in the README.
…v1 ledger AddDeviceRules(path, rules) becomes SetDeviceRules(handle, rules): make the cgroup's daemon-owned grants equal exactly rules. The processor opens the cgroup directory once (O_DIRECTORY|O_RDONLY|O_CLOEXEC) and records its path and inode from Fstat on that descriptor; v2 queries and attaches against it and v1 reaches devices.list, devices.allow and devices.deny with openat, so a cgroup recreated at the same path cannot receive the old container's rules. cgroup v1 gets a baseline-aware ledger, owned by the Processor and keyed by path plus inode. Set computes the runtime baseline as devices.list minus the owned bits, owns only desired bits outside it, revokes before it grants (stopping on a failed deny, so A to B never ends as A plus B), updates the ledger after every successful write, restores owned bits removed externally, and re-reads devices.list to report drift as a retryable error. Baseline bits are never denied and redundant grants never recorded. cgroup v2 keeps its grant-only wrapping behind the new signature for now. README and SECURITY.md document the v1-only gap: grants made by a previous daemon instance look like baseline after a restart and are not revoked.
…owned blocks Every program the daemon attaches is now HEADER, init, rule blocks, TRAILER, original. Header and trailer are dead stores to R0 carrying a magic, a format version, a per-lineage 64-bit nonce, a 128-bit SHA-256 digest of the original and the block length. A program counts as owned only if the whole wrapper regenerates byte for byte from the rules and nonce parsed out of it; anything that starts like a header but fails is ErrOwnedBlockConflict and nothing is changed (reason owned_block_conflict, ERROR with the cgroup and the remedy: restart the container). Unmarked programs are never interpreted. SetDeviceRules strips validated wrappers back to the original before wrapping again, so grants replace each other instead of stacking, with no daemon state. Empty rules restore the bare original under an owned wrapper and leave unmarked programs alone. Wrappers are named sda_devfilter and are swapped in through the pairwise swap primitive; every load happens before any attach. A reloadability gate limits wrapping to the instruction subset device filters use (context loads at 0/4/8, mov/and/or/lsh/rsh, in-range conditional jumps and ja, exit; no calls, maps, 64-bit immediates, stores, atomics or may_goto) with map IDs known and empty on Linux 4.16+. Golden fixtures of real runc 1.5, crun 1.28 and systemd 259 device filters captured on Linux 7.0 prove they pass the gate, reload unchanged, and that the kernel's dump of the wrapper strips back to the exact original; a test-only interpreter checks the verdicts. BREAKING CHANGE: device filters outside the wrappable instruction subset (or on kernels older than 4.16) are no longer wrapped; the container is retried with reason program_not_wrappable. Upgrade note: grants attached by earlier versions carry no marker and look like the runtime's own filter, so they persist until the container restarts; restart containers processed by earlier versions.
Attached programs are classified into wrapper lineages (validated nonce plus byte-identical innermost original) and bare runtime programs. With grants, each lineage ends as one wrapper equal to the desired one (an equal member is kept, otherwise the first is replaced) and its other members are detached; each bare program is wrapped as its own pair under a new nonce, so identical runtime programs keep their count and the leftover of a failed rollback ends as one redundant, identical wrapper. Without grants, lineages go back to their bare original, and a lineage whose original is already attached bare only has its wrappers detached, which is how an interrupted restore converges. A wrapper over a wrapper is re-emitted as one. Kernel program IDs are never used, every comparison is on canonical bytes, and every attach precedes the detach it enables. A per-cgroup FilterCache (process memory, keyed by path plus inode, with Forget for lifecycle ends) records the runtime originals and lineage nonces. When systemd's daemon-reload has detached every filter, the cached originals are attached again, bare or wrapped under their nonce, restoring the runtime's own restrictions as well; a rebuild that stops halfway stays marked in the cache so the next pass reattaches the lineages still missing instead of forgetting them. With nothing cached, ErrFilterMissing: the processor treats a privileged container as a no-op (INFO, skipped as privileged_no_filter) and anything else as retryable with reason filter_missing and a warning to restart the container, never a silent success.
…dentity pinning Processor.Reconcile (formerly ProcessContainer) computes the complete device set the current config allows a container and always applies it, including the empty set: a container disabled or not opted in by policy, one with invalid labels and one with any unresolved device all get the empty set, so an earlier grant is revoked instead of kept. An incomplete device set returns a retryable incomplete_device_set error. The per-container mutex is held from the config load through SetDeviceRules. The container process is pinned with a pidfd, /proc/<pid>/cgroup and mountinfo are read through one directory descriptor (cgroup.ParseProcCgroup), and immediately before the mutation the pinned process must be alive, a second bounded inspect must report the same pid, StartedAt and running state, and, for grants only, the pid must be listed in cgroup.procs read through the cgroup handle (CgroupHandle.HasProcess). The verified identity is recorded per lifecycle (container ID plus StartedAt) before any mutation and kept across failed applies. A container that is no longer running, or that Docker cannot inspect, has the empty set applied in the cgroup its lifecycle was last verified in, only if the re-opened path still has the recorded inode; a removed or recreated cgroup releases the entry without mutation (Ledger.Forget, FilterCache.Forget). An inspect failure keeps the container pending; one caused by the caller cancelling (daemon shutdown) revokes nothing, while a timeout does. Entries of exited lifecycles are retained until their cgroup is verified gone; the periodic sweep that does so lands with the lifecycle records. Dry-run pins nothing, reads no /proc entry and builds no cgroup API; it logs the rules and a would-set line per container. README documents the reconcile semantics, the revoke-then-regrant cost of transient errors and the Linux 5.3 pidfd requirement.
A single coordinator goroutine now owns reconciliation passes: it tracks the latest config generation and a trigger epoch bumped by every request, lists containers with a bounded ContainerList (retrying the whole pass with minBackoff/maxBackoff when the list fails), retries failed containers from a pending set with per-container backoff until they succeed, disappear (Processor reports ErrContainerGone) or the daemon stops, and owns the processed dedup map (with its TTL) that the event consumer reaches only through mutex-guarded methods. A newer request supersedes an in-flight pass after its current container, re-tags pending entries and coalesces into one trailing pass; results from an older generation or epoch never clear the gauges or log completion. Run waits for the coordinator before returning. New gauges sda_reconcile_pending_containers and sda_reload_incomplete; the daemon warns while either is non-zero and logs config reload complete once both are zero for the latest request. config.NewStore now returns the Store and a Publisher, and every publication bumps a generation. Processor.PublishAndReconcile publishes under the processor lock and hands the generation to the coordinator; it is the only publication path after startup (enforced by a source test). SIGHUP (now registered before startup work), systemd reloads and startup all request passes; the startup pass reconciles every running container, so a live start replaces or strips grants a previous instance left, while a dry-run start warns that it cannot. README, SECURITY.md and the architecture metrics table document the passes, the gauges and that restarting into dry-run is not a cleanup path; a new enforcement test (SDA_IT_ENFORCE=1) covers the kill-and-restart cases.
…erified gone The coordinator now keeps a per-container history of runs (processor.Lifecycles, installed into the processor, which records into it): provisional when an inspect failed, verified once the run's start time and cgroup identity were verified against the pinned process, terminal once the run ended. A restart under the same ID adds a record, so late cleanup for an old run never reaches the new one. The processor's per-lifecycle identity map is replaced by these records. die and destroy are added to the event subscription. The coordinator adjudicates every event: the consumer only submits it. A die or destroy resolves to the verified run with the latest start not after the event time (else the provisional records), whose cgroup gets the empty set; the record stays until its cgroup is verified gone. Before handling a container the coordinator reserves it, so work requested meanwhile (an event, a pass, a cleanup) coalesces into one follow-up; a start event not newer than a running reconcile is covered by it, while a cleanup covers none. Pending retries skip reserved containers, and completion waits for running reconciles. Records, with the cgroup v1 ledger entry and the cached cgroup v2 filters of their cgroup, are released only when the cgroup path no longer exists or its inode changed; an inspect not-found never releases one. A periodic sweep re-checks every record, retries revokes of ended runs that failed, and drops records that never named a cgroup once ended or stale.
… the config file config.LoadFile now makes two passes over the same bytes. The first walks the parsed nodes, following aliases, and requires every known key to have exactly its shape: dry-run and log-time a YAML bool (a quoted yes is no longer coerced), policy-mode, log-format, log-level and docker-socket a non-empty string, metrics-addr and debug-addr a string that may be empty, device-allow and device-deny a list of non-empty strings. An explicit null, also behind an alias, is an error instead of silently meaning not set, and [null] can no longer decode to [""], which read as allow-all. Merge keys are refused. The second pass decodes with unknown keys rejected (a device_deny typo no longer drops the deny list) and requires a single document. log-format, log-level and policy-mode are validated when the file is loaded and again, on the effective values, at startup and before a SIGHUP reload is applied; the undocumented log-level aliases warning, fatal and panic are retired (README flag section). The README documents the strict file rules, and the trust-boundary checklist is updated.
…k to the startup file Right after flag.Parse the flag values and the set of flags given on the command line are snapshotted into one settings struct (list flags deep-copied). One pure mergeSettings(flags, cliSet, file) returns the complete effective settings and replaces the duplicated merges of applyFileConfig and watchSIGHUP, which wrote the startup file's values into the flag globals so that a key later removed from the file kept its old file value (policy-mode: all stayed all). Startup consumes every field. A SIGHUP reload validates the merged settings (enums and policy) before anything is applied, applies only the hot-reloadable ones (logging, dry-run, policy), logs and ignores a changed docker-socket, metrics-addr or debug-addr, and publishes through Processor.PublishAndReconcile. A reload may turn dry-run off but never on: that is rejected with dry-run cannot be enabled on a running daemon, and the previous config is kept. Any rejected reload leaves the store, the logger and the in-force settings untouched. SIGHUP is registered in run() and the channel passed in. README (Config File, -dry-run row) and the trust-boundary checklist describe the reload rules; config_test.go replaces the empty main_test.go.
…VNAME
Every name under a /dev mount (the source, walked entries and symlinks) now goes through one evaluation on file descriptors: a lexical gate (clean, under /dev), openat2 beneath a /dev descriptor opened once per pass with RESOLVE_BENEATH and RESOLVE_NO_MAGICLINKS, a bounded by-hand follow of absolute symlinks back into /dev (EXDEV), fstat for type and major:minor, and the canonical name from the device's sysfs uevent DEVNAME. Names that lead outside /dev are skipped with a WARN and counted in sda_device_candidates_skipped_total{reason}.
Policy splits DeviceAllowed into Denied, checked on the alias, the resolved node and the canonical name, and Authorized, checked on the canonical and resolved names. Candidates are aggregated per device across all mounts: a device is granted only when authorized with no deny vote under any name, replacing the first-wins dedup. A candidate whose identity cannot be established and that policy would otherwise grant empties the whole desired set with a retryable error.
The daemon refuses to start without openat2 RESOLVE_BENEATH; the documented kernel requirement is now Linux 5.6.
…d completely A directory mount is walked only to enumerate names, and every name goes through evaluateCandidate. A root that cannot be opened, a directory that cannot be read, any walk error, or more than 4096 entries in one mount (mount too large; narrow the bind mount) now makes the container's whole desired set empty with a retryable error, like an unresolved candidate: an alias that was not seen might deny a device that was. The silent depth cap is gone and there is no subtree exception. The enforcement test now creates, inside the opted-in container, a node with the granted device's major and the next minor and one of the other device type, and requires both to stay denied. README, SECURITY.md and docs/architecture.md state that the canonical identity is the security authority, that alias-level deny is a best-effort veto, recommend deny globs on kernel names such as /dev/sd* and /dev/nvme*, and explain how to see a device that lacks DEVNAME and why whole-/dev mounts can hit the entry cap. The trust-boundary skill records the containment and planted-node gaps as closed.
Add a test job to the CI workflow with the same steps as the release workflow's test job: setup-go pinned by SHA with Go from go.mod, then make go-vet and make go-test. Until now vet and unit tests only ran at release time.
In every example compose file the commented daemon flags (-device-allow, -policy-mode) sat before the image in the docker run argument list, so uncommenting one made it a docker run flag instead of reaching the daemon. They now follow the image, as in deployments/docker/docker-compose.yaml, with a note saying why.
USB device nodes live at /dev/bus/usb/<bus>/<device>, and a glob * does not cross a slash, so /dev/bus/usb/* matched only the bus directories and granted no device. The flag comment and the label now use /dev/bus/usb/*/*.
The -device-allow and -device-deny flags take one glob each and are not split on commas, so the smartctl-exporter example's comma list would have been one glob that matches nothing. It now repeats the flag per glob, and the README flag table says one glob per flag, unlike the comma-separated labels.
The gpu-passthrough and smartctl-exporter examples mounted the host DBus socket at /run/dbus/system_bus_socket inside the daemon container, but the daemon connects to /var/run/dbus/system_bus_socket and dhi.io/static has no /var/run -> /run symlink, so reload handling stayed disabled. Both now mount it at /var/run/dbus/system_bus_socket, as the reference compose file does.
…to it The Swarm wrapper ran a bare docker run with no --name, so the README's SIGHUP recipe could not find the daemon, and a killed wrapper orphaned the daemon while Swarm started a second one. The wrapper is now an sh -c script in the README, deployments/docker/docker-compose.yaml and all five examples. It sets a TERM and INT trap first, removes a stale swarm-device-access container, starts docker run --name=swarm-device-access in the background (a shell blocked on a foreground child defers its traps), and waits on it; the trap runs docker stop -t 10 on the daemon, waits for it and exits with its status. stop_grace_period is 30s so the stop can finish. The docker run arguments are built with one set -- line each, options before the image and daemon flags after it, so an optional mount or flag is enabled by uncommenting one line and cannot land on the wrong side of the image. The SIGHUP recipe becomes docker kill -s HUP swarm-device-access, and the commented config example mounts the directory, since a single-file mount keeps the old content after an editor replaces the file. docs/testing.md adds the manual Swarm check: kill -9 the wrapper task and confirm the replacement starts without a name conflict and exactly one daemon container runs. That check needs a Swarm host and was not run here; the script's signal handling and argument building were checked with a stub docker under busybox sh.
…e code AGENTS.md, docs/architecture.md and the tracked .serena/memories/core.md described an event loop in main.go with processExistingContainers, processContainer and AddDeviceRules, none of which exist. They now describe internal/daemon (Run, listenEvents, consumeEvents, processOne and the coordinator that owns passes, retries, reservations and the time-stamped processed dedup with processedTTL), internal/processor (Reconcile, CollectMountRules, applyPinned and SetDeviceRules through a cgroup handle), the config, policy, daemon, processor and observability packages, and the current v1 and v2 file names, with the reconcile path described once. CONTRIBUTING.md drops the link to the missing UPSTREAM_AUDIT.md and says go-build targets the host's GOARCH and docker-build is single-arch. AGENTS.md and CONTRIBUTING.md say release notes come from gh release create --generate-notes (pull request titles), not commit messages. SECURITY.md notes that Swarm supports cap_add since Engine 20.10 and why the wrapper is still needed. Metrics and debug examples bind 127.0.0.1, and log examples use the current setting device rule message.
Any client on the system bus can emit a signal named org.freedesktop.systemd1.Manager.Reloading, and the watcher counted every one. The match rules now name systemd as sender and /org/freedesktop/systemd1 as path, and each signal is also checked against the unique name that owns org.freedesktop.systemd1, resolved with GetNameOwner and updated only from the bus daemon's NameOwnerChanged, so a systemd re-exec does not silence the watcher and a peer cannot redirect it. The signal channel is registered before the owner is looked up, so an owner change that races the lookup is queued instead of dropped. Re-applies are coalesced: reloads that complete while one runs are absorbed into one trailing run after it. Tests cover foreign senders, foreign object paths, spoofed owner changes, following a real owner change, and coalescing.
A re-subscription only replays the events dockerd still buffers, and a restarted dockerd has none, so containers started during the gap could go unreconciled. Every successful re-subscription now requests a coordinator pass that re-lists and reconciles every running container; the coordinator's processed map skips replayed start events the pass already covered. A new test drives a reconnect with no explicit request and events arriving on the new stream while the pass runs, under -race, and checks that the pass lists and reconciles both running containers.
A glob that is relative, not clean or not under /dev/ can never match a device path, so a deny written that way silently denied nothing. ValidateGlobs now rejects such patterns, so flags and the config file fail at startup or reload and a label makes the container fail closed. The README explains that * does not cross a slash, that /dev/bus/usb/* does not match /dev/bus/usb/001/002, suggests /dev/bus/usb/*/*, and states the new pattern rules. Tests cover relative, uncleaned and outside-/dev patterns in ValidateGlobs, labels and the global policy.
The CI and integration workflows ran every job with the repository's default token permissions. Both now default to no permissions at the top level, and each job states its own: the reviewdog jobs (actionlint, pre-commit-hooks, markdownlint, shellcheck, yamllint) get contents read and pull-requests write, since every one of those actions reports with -reporter=github-pr-review and none reports checks, so checks write is not granted. audit-deps, test and integration get contents read only. conventional-commits also gets contents read only: checked against the action's source at the pinned SHA, it takes no token and only prints annotations, so pull-requests write would be unused.
…e run Fork pull requests get no secrets, so their integration job failed at the dhi.io login. The job now checks whether a pull request's head is another repository (head.repo.full_name differs from github.repository; head.repo.fork alone is also true for this repository's own branches if the repository is itself a GitHub fork): such a pull request skips the login and runs only the dry-run tests, under the job name integration (fork: enforcement skipped). Same-repo pull requests and pushes keep the login and SDA_IT_ENFORCE=1 mandatory, and a missing DHI_TOKEN there fails the job; the gate is never whether the secret is set, which would silently downgrade trusted runs. The release workflow's integration job now mirrors it: dhi.io login, SDA_IT_ENFORCE=1 and SDA_IT_REQUIRE_RELOAD=1, and the SKIP guard. docs/release.md and docs/testing.md describe both. Verifying both branches needs GitHub runs (one same-repo pull request, one fork pull request) and was not done here; actionlint and the YAML checks pass.
The syntax directive pulled docker/dockerfile:1.7 by tag only, so a rebuild of the same commit could use a different frontend; it is now pinned by its index digest like the base images. The deps stage downloaded modules but nothing used it, since build started again from the golang base; build now starts FROM deps, so the module download runs before the source is copied. Verified with make docker-build.
The metrics and debug servers bound their address in a goroutine, so a busy or invalid address was only an error log while the daemon ran on without them. StartMetricsServer and StartDebugServer now bind synchronously and return a stop func or the error, then serve on the bound listener in a goroutine. run() exits 1 when either fails, and stops the metrics server again if the debug server then fails. The metrics server gets a 10s read, 30s write and 60s idle timeout. The debug server gets 10s read and 60s idle but no write timeout, because a CPU profile or trace streams for as long as the caller asks; the README and docs/architecture.md say it must stay on a loopback address. Tests cover a busy address, /healthz and /readyz transitions, /metrics, stop and context shutdown, a profile?seconds=1 request completing, and a failing second server releasing the first one's address.
The mount-excluded WARN built its allow_globs and deny_globs fields with append on the global policy's slices, which the config store shares with every reconcile; with spare capacity, append wrote the container's globs into the stored slice's backing array. The fields now use slices.Concat. A test gives both global lists spare capacity and checks that the backing arrays are untouched after the WARN.
The event loop, the coordinator, the systemd watcher, the observability servers and run() captured logger.L() once at start, so a SIGHUP that changed the log level or format never reached them. They now call logger.L() at each log site. A test changes the logger while the systemd watcher runs and checks that the next line goes to the new one.
A mount was walked whenever its Source was under /dev, whatever its type, but only a bind mount's Source is a host path. The processor now requires Type bind before treating a mount as a device mount; volumes, tmpfs and other types are ignored, and /dev is not even opened for a container without a /dev bind mount. A test covers volume, tmpfs, named pipe and untyped mounts with a /dev Source next to a real bind mount.
The consumeEvents tests in events_test.go and daemon_test.go ended by cancelling 20ms after starting, which either raced a slow runner or waited for nothing. A test expecting N applies now cancels from inside apply once the N-th call is observed (cancelAfter), never earlier. A test expecting zero applies runs the consumer in a goroutine and waits for a positive signal that the event was handled, the coordinator's new debug line event already covered by a pass; skipped, before it cancels, with a two-second wait kept only as the documented upper bound.
…rtup The watcher now runs on its own goroutine, started before the startup pass, so a slow or wedged system bus never delays the pass or the event loop. Each connection attempt is bounded, a failed attempt or a lost subscription (a DBus restart) is retried with backoff, and every successful subscription requests a full pass, since a reload that completed while no watcher was subscribed, the startup pass included, was never seen.
A running container can move to another cgroup under the same run, for example when systemd inside it moves its processes into a child cgroup. The record used to be overwritten with the new cgroup, leaving the grants on the previous one, where they keep applying to every descendant. The previous cgroup now gets the empty set first; until that succeeds the new cgroup is not granted and the record keeps naming the previous one, so a retry revokes it again.
A daemon now requests one pass at startup and another when its reload watcher subscribes, and the second can land after a test has started counting processed records. waitReady now also waits for the watcher's outcome and for the last of those passes to complete, so every later record comes from what the test does. The reload subtest also explains why the restarted daemon finds a filter to wrap.
Revoking the previous cgroup is real work between the pre-mutation check and the grant, and the run can move again meanwhile. The pinned process is now verified once more after the revoke, cgroup membership included, before the new cgroup is recorded or granted.
A setup attempt that timed out used to keep running in the background with its connection open, and on a wedged bus every retry added another. The connection is now bound to a per-attempt context that the setup timeout cancels: godbus closes it, which fails the handshake or call still waiting, and the attempt returns before the next one starts. The match rules and the owner lookup take the same context.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the correctness, security and drift issues found in a full audit of the project. There are 32 commits, one per item, and each one builds, passes the tests and passes
make check-stageby itself, so reading them one at a time is the easiest way to review this.Highlights
Processor.Reconcile. It sets a container's grants to exactly what the current policy allows, including none, so anything uncertain revokes. One coordinator owns passes and retries, and there are new gaugessda_reconcile_pending_containersandsda_reload_incomplete.BPF_F_REPLACEwhere the kernel supports it, otherwise the new filter is attached before the old one is detached. Grants live in validated owned wrappers (sda_devfilter), so they replace each other instead of stacking and survive a daemon restart. Filters wiped bysystemctl daemon-reloadare rebuilt.dieanddestroy, and keeps a lifecycle history until each run's cgroup is verified gone./devwithopenat2(RESOLVE_BENEATH)and identified byfstatplus the sysfsDEVNAME. Deny globs are checked on every name a device is found under. Planted nodes and paths that escape/devno longer bypass policy, and an unknown identity or an incomplete walk empties the set./dev/. systemdReloadingsignals are only accepted from systemd itself. After an event-stream reconnect the daemon re-lists and reconciles every running container. A busy metrics or debug address stops startup.Breaking changes
deploy.labelsare ignored. Putswarm-device-access.*labels under the service's top-levellabels:.BPF_F_ALLOW_MULTI, one the daemon cannot read, or one outside the wrappable instruction subset makes the daemon skip or retry the container instead of making a best-effort change.openat2); on older kernels the daemon refuses to start./dev/are rejected.Not verified here
SDA_IT_ENFORCE=1) were not run locally; CI runs them on this PR. Locally only the dry-run integration suite ran.kill -9the wrapper task, then confirm the replacement starts cleanly with exactly one daemon container) needs a Swarm host.