Skip to content

About

Cron-scheduled Docker host vulnerability scanner using Trivy — structured JSON logs for OpenTelemetry pipelines.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

runtime-vul-detect

Checks Verify Release Release Test Coverage (Python) Test Coverage (Go) Trivy HIGH Trivy CRITICAL License

Go Python Docker

A lightweight, self-hosted Docker host vulnerability scanner. It discovers every image in use across all containers on the host, scans them with Trivy on a cron schedule, and emits structured JSON logs — ready to be collected by Grafana Alloy, Promtail, or any log shipper that feeds into Loki or a similar backend.

What is a vulnerability?

A software vulnerability is a weakness or flaw in a program that can be exploited to compromise a system. The industry tracks them through a set of well-known standards:

Term Meaning
CVE Common Vulnerabilities and Exposures — a unique identifier (e.g. CVE-2021-44228) assigned to each known flaw. Catalogued by MITRE.
NVD National Vulnerability Database — NIST's enriched feed of CVEs with CVSS scores, affected package ranges, and patch availability.
CVSS Common Vulnerability Scoring System — a 0–10 score that reflects exploitability and impact. Scores ≥ 9.0 are CRITICAL, 7.0–8.9 are HIGH.
SCA Software Composition Analysis — scanning a software artifact (container image, lock file) against known CVE databases to surface vulnerable dependencies.

Trivy performs SCA against the OS package layer and language-ecosystem packages (pip, npm, gem, etc.) inside each container image, matching installed versions against the NVD and vendor advisories.

Why runtime-vul-detect?

Container images accumulate CVEs over time. A base image that was clean at build time can become vulnerable days later when a new CVE is published against one of its packages. Most teams only scan at CI time — by the time a vulnerability is disclosed, the image is already in production and nothing re-checks it.

Market landscape

Tool Self-hosted Standalone Docker Scheduled Structured logs Notes
runtime-vul-detect Yes Yes Yes JSON This project
Trivy CLI Yes Yes No JSON Excellent scanner; no scheduler, no deduplication, no log shipping
Grype Yes Yes No JSON Strong alternative scanner (Go binary); same gap as Trivy CLI
Trivy Operator Yes No — k8s only Yes CRD events Kubernetes CRD-based; requires a cluster
Kubescape Yes No — k8s only Yes – Uses Grype under the hood; cluster-coupled
Docker Scout No (SaaS) Yes No – Cloud-only; limited free tier
Snyk Container No (SaaS) Yes No – Phones home to Snyk cloud for analysis
Clair Yes Yes No API Registry-first microservice; requires Postgres; no CLI
Aqua Security Yes (paid) Yes Yes – Commercial platform built on Trivy; heavyweight
Falco Yes Yes – JSON Runtime behavioural detection, not CVE scanning

The gap: For a standalone Docker host (no Kubernetes), there is no lightweight self-hosted tool that combines scheduling, image deduplication, and structured log output out of the box. runtime-vul-detect closes this gap. Future Kubernetes support is on the roadmap, which would put it alongside Trivy Operator but with a simpler operational model (no CRDs).

Architecture

┌──────────────────────────────────────────────────────────────────────────┐
│  Docker host                                                             │
│                                                                          │
│  ┌──────────────────────┐  tcp:2375   ┌──────────────────────────────┐  │
│  │  runtime-vul-detect  │◄────────────│  docker-socket-proxy         │  │
│  │                      │  GET-only   │  wollomatic/socket-proxy:1.12.1   │  │
│  │  APScheduler         │  allowlist  │                              │  │
│  │    ├── scan (cron)   │             │  UID 65534 (nobody:nobody)   │  │
│  │    └── db-update     │             │  read-only rootfs            │  │
│  │          │           │             │  ALL capabilities dropped    │  │
│  │          ▼           │             └──────────────┬───────────────┘  │
│  │     Trivy binary     │                            │                  │
│  │          │           │              /var/run/docker.sock (ro)        │
│  │          ▼           │                                               │
│  │   JSON log lines ────┼──► stdout                                     │
│  └──────────────────────┘        │                                      │
│                                  ▼                                       │
│                         Grafana Alloy / Promtail                         │
│                                  │                                       │
│                                  ▼                                       │
│                             Loki / SIEM                                  │
└──────────────────────────────────────────────────────────────────────────┘

Docker socket proxy

The scanner never mounts /var/run/docker.sock directly. Instead, a wollomatic/socket-proxy sidecar holds the real socket and exposes a TCP endpoint on an internal-only Docker bridge network (socket-proxy, internal: true).

Why this matters: The Docker socket is effectively root on the host — any process that can write to it can start privileged containers, mount the host filesystem, or escape to the host entirely. A compromised scanner (e.g. via a malicious Trivy DB or a CVE in a dependency) would otherwise have full daemon access. The proxy eliminates that path.

What the proxy enforces:

Control Detail
Protocol TCP only; no Unix socket exposure to the scanner
Methods GET only — no POST, DELETE, PUT
Endpoints containers/json, images/sha256:.../json, version, /_ping — everything else returns 403
Network internal: true bridge — the proxy container has no internet route
Privileges UID 65534 (nobody), all Linux capabilities dropped, read-only rootfs

Even if the scanner container is fully compromised, the attacker can only enumerate container names and image SHAs — they cannot start, stop, exec into, or delete anything on the host.

Resource usage

Resource At idle During scan Notes
Memory (RSS) ~60–80 MB ~300–500 MB peak Python interpreter + libraries idle; Trivy loads its DB (~250 MB) per invocation
CPU ~0% 100% single core (5–30 s per image) Trivy is a Go binary — scans are CPU-efficient; Python orchestrator is idle otherwise
Disk (image) ~250 MB – Python 3.13-alpine + Trivy binary + Cython .so files
Disk (cache volume) ~300–500 MB – Trivy vulnerability DB + layer cache; grows slowly as DB is updated
Network 0 ~50–100 MB on DB update Incremental DB download from GitHub; image scans use local Docker socket only
Docker socket Not mounted Not mounted Access goes through the socket proxy sidecar; scanner never touches the socket directly
Socket proxy ~2–4 MB ~2–4 MB wollomatic/socket-proxy:1.12.1; distroless, UID 65534, all caps dropped; 64 MB memory limit

Cython impact: Compiling .py → .so reduces container startup time slightly (no bytecode compilation at import) but does not materially reduce steady-state memory because the Python interpreter and its C extension infrastructure are still present. The primary benefit is source code protection.

Trivy subprocess cost: Each scan spawns a trivy image subprocess. Trivy loads its vulnerability DB on every invocation (~100 MB read from disk). For hosts running many unique images, a persistent Trivy server mode (trivy server) would amortize this cost — planned as a future optimisation.

Quick start

git clone <repo>
cd runtime-vul-detect
chmod +x *.sh && ./up.sh python   # or: ./up.sh go

All orchestration scripts live at the repo root (so the python/ and go/ folders stay pure code + Docker assets). ./up.sh <variant> builds the image, starts the container, waits for the Docker healthcheck to pass, then tails logs. Stop it with ./down.sh python (or go) — add --purge to drop the cache volumes. up.sh and down.sh act on a single variant; build.sh, scan_vul.sh, test_up.sh and verify_metrics.sh also accept all.

The Trivy version is pinned in <variant>/docker-compose.yml under build.args.TRIVY_VERSION. Update it there to upgrade. Run ./build.sh python (or go / all) to rebuild and immediately scan the produced image for vulnerabilities.

Implementations: Python (main) & Go

The project ships two interchangeable implementations of the same scanner. They share identical behaviour, configuration (python/config/config.yaml + python/config/ignore.yaml), and JSON log schema — pick whichever fits your operational constraints.

Python (main / reference) Go (alternative)
Source python/ — src/, tests/ go/ — cmd/, internal/
Base / runtime python:3.13-alpine, Cython-compiled .so gcr.io/distroless/static-debian12:nonroot, single static binary
Scheduler APScheduler robfig/cron/v3 (SkipIfStillRunning)
Docker access docker SDK Go stdlib net/http over the unix socket (no SDK, no containerd/otel deps)
Logging python-json-logger log/slog (stdlib)
Published tag ghcr.io/jwang996/runtime-vul-scanner:1.0.0-python-alpine ghcr.io/jwang996/runtime-vul-scanner:1.0.0-go-distroless
Image size ~250 MB ~30–50 MB
Idle memory ~60–80 MB ~8–15 MB

Python remains the main implementation — it is the reference for behaviour and documentation. The Go port is a drop-in alternative for environments that want a minimal distroless image, a smaller attack surface (no shell, no package manager, no Docker SDK dependency tree), and lower memory/startup cost.

Drive either implementation with the root scripts — the variant folders hold only source + Docker assets:

chmod +x *.sh
./up.sh python           # build, start, wait for healthy, tail logs (one variant)
./build.sh python        # rebuild + Trivy-scan the produced image (or: go / all)
./test_up.sh all         # containerised unit tests + coverage gate (both variants)
./verify_metrics.sh all  # bring up with all /metrics protections and assert them
./verify_webhook.sh all  # bring up against the vulnerable target and assert webhook alerts
./down.sh python         # stop the variant (add --purge to drop volumes)

up.sh and down.sh act on a single variant (python or go) — the two share container names and bind port 9090, so running both at once is neither needed nor supported. build.sh, scan_vul.sh, test_up.sh and verify_metrics.sh accept python | go | all.

Both implementations are built, scanned (Image Scan workflow), unit-tested and published (CI workflow) independently in CI. Test files in the Go version are co-located _test.go files (Go convention) and are never compiled into the runtime binary.

Configuration

All scanner behaviour is controlled by two YAML files mounted read-only into /config/ inside the container. Both files have sensible values out of the box.

Hot-reload: the scanner polls the config and ignore files every 15 s and applies changes without a container restart. Editing scan_cron, db_update_cron, severity, ignore_files, db_update, or timezone reschedules the running jobs in place (a Config reloaded line is logged). Invalid new config — unparseable YAML or a bad cron/timezone — is logged as an error and the previous config keeps running. scan_on_start, the metrics_* settings, the webhook_* settings, and the events_* settings are applied at startup only; changing those requires a restart.

python/config/config.yaml — scanner settings

# IANA timezone for cron schedule interpretation
# Full list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones
timezone: "Europe/Berlin"

# How often to scan all images (standard 5-field cron, in the timezone above)
scan_cron: "*/5 * * * *"

# How often to refresh the Trivy vulnerability database
db_update_cron: "0 0 * * *"

# Severities to include in results — anything below this threshold is silently dropped
# Options: CRITICAL, HIGH, MEDIUM, LOW, UNKNOWN
severity: "CRITICAL,HIGH"

# Keep the Trivy CVE database up to date automatically
db_update: true

# Run a full scan immediately when the container starts (before the first cron tick)
scan_on_start: true

# Trivy cache directory — must match the volume mount in docker-compose.yml
cache_dir: "/tmp/trivy-cache"

# Ignore files to load at startup — multiple files are merged; duplicates are logged as errors
ignore_files:
  - "/config/ignore.yaml"

# Prometheus /metrics endpoint (severity counts as gauges + scan meta)
# Applied at startup; changing these requires a container restart
metrics_enabled: true
metrics_port: 9090
Key Default Description
timezone Europe/Berlin IANA timezone string used to interpret both cron expressions
scan_cron 0 * * * * How often to scan all images (hourly by default)
db_update_cron 0 0 * * * How often to pull the latest CVE database (nightly by default)
severity CRITICAL,HIGH Which severity levels to include; lower levels are filtered before logging
db_update true Toggle automatic database updates; disable when airgapped
scan_on_start true Scan immediately on startup rather than waiting for the first cron tick
cache_dir /tmp/trivy-cache Trivy's local database and layer cache
ignore_files [/config/ignore.yaml] List of ignore config paths; all are merged at startup
metrics_enabled true Serve the Prometheus /metrics endpoint; set false to disable the HTTP server
metrics_port 9090 Port the /metrics endpoint listens on inside the container
metrics_basic_auth_user "" Require HTTP Basic auth on /metrics when set (together with the password)
metrics_basic_auth_password "" Password for Basic auth
metrics_tls_cert_file "" PEM cert path; with the key, serves /metrics over HTTPS
metrics_tls_key_file "" PEM private-key path for the TLS cert
metrics_tls_client_ca_file "" CA bundle to require & verify client certificates (mutual TLS)
webhook_enabled false Opt-in master switch for webhook alerting on newly-appeared findings; false = no HTTP at all (pure monitoring)
webhook_url "" HTTP(S) endpoint that receives the POSTs
webhook_timeout_seconds 10 Per-request HTTP timeout in integer seconds
webhook_auth_header "" Optional HTTP header name (e.g. Authorization) to attach to each POST
webhook_auth_token "" Optional secret value; sent as <webhook_auth_header>: <webhook_auth_token> only when both are non-empty (never logged)
events_enabled false Opt-in master switch for event-driven scanning via the Docker /events stream; false = no /events connection at all (pure cron behaviour)
events_debounce_seconds 5 Idle window over which event image refs are coalesced before one targeted scan fires per distinct ref

Note: LOG_LEVEL is the only setting not in config.yaml. It must be set as a LOG_LEVEL environment variable because the logger initialises before the config file is read. Default is INFO.

python/config/ignore.yaml — suppressing known non-issues

Use this file to whitelist CVEs that are not exploitable in your environment, or to skip scanning specific images entirely.

# Skip these images entirely — supports fnmatch wildcards
images:
  # - "busybox:*"           # skip all busybox tags
  # - "internal-tool:dev"   # skip a specific dev image

# Suppress specific CVEs — uses Trivy's native .trivyignore.yaml format
# See: https://aquasecurity.github.io/trivy/latest/docs/configuration/filtering/
vulnerabilities:
  # - id: CVE-2021-44228
  #   statement: "log4j not reachable — no Java runtime in this image"
  #   expired_at: 2027-01-01     # optional: unquoted YAML date; re-surfaces after this date
  #
  # - id: CVE-2023-1234
  #   paths:
  #     - usr/local/lib/python3.13/site-packages/requests
  #   statement: "requests not exposed to untrusted input"

You can supply multiple ignore files by listing them under ignore_files in config.yaml. If the same CVE ID or image pattern appears in more than one file, an ERROR log is emitted at startup — this makes cross-file conflicts visible immediately rather than silently winning in undefined order.

Log output

Every scan result is emitted as a single JSON line on stdout:

{
  "timestamp": "2026-05-28T12:00:00Z",
  "level": "info",
  "logger": "runtime-vul-detect",
  "message": "scan_result",
  "event": "scan_result",
  "image": "nginx:1.21",
  "image_id": "sha256:abc123...",
  "containers": ["a1b2c3d4e5f6", "b2c3d4e5f6a1"],
  "total_vulnerabilities": 3,
  "severity_counts": {
    "CRITICAL": 1,
    "HIGH": 2,
    "MEDIUM": 0,
    "LOW": 0,
    "UNKNOWN": 0
  },
  "vulnerabilities": [
    {
      "id": "CVE-2021-1234",
      "package": "openssl",
      "version": "1.1.1k-r0",
      "fixed_version": "1.1.1l-r0",
      "severity": "CRITICAL",
      "title": "Memory corruption in openssl",
      "target": "alpine:3.15"
    }
  ]
}

containers lists all container short IDs sharing that image, so you can trace a finding back to the exact running workload.

Collecting with Grafana Alloy

// Discover the scanner container and forward its stdout to Loki
loki.source.docker "scanner" {
  host       = "unix:///var/run/docker.sock"
  targets    = discovery.docker.containers.targets
  forward_to = [loki.process.parse_json.receiver]
  labels     = { job = "vulnerability-scanner" }
}

// Parse the JSON and promote scan fields as structured metadata
loki.process "parse_json" {
  stage.json {
    expressions = {
      event       = "event",
      image       = "image",
      total_vulns = "total_vulnerabilities",
      critical    = "severity_counts.CRITICAL",
    }
  }
  stage.labels {
    values = { image = "", event = "" }
  }
  forward_to = [loki.write.default.receiver]
}

Because each line is valid JSON, Loki's json parser can extract severity_counts, image, and total_vulnerabilities as structured fields or metric values without regex.

Prometheus metrics

Alongside the JSON logs, the scanner serves a Prometheus exposition endpoint at GET /metrics (enabled by default; see metrics_enabled / metrics_port). It is published on the host loopback only — 127.0.0.1:9090 in docker-compose.yml — so it is not exposed on the public interface. Both the Python and Go implementations emit byte-identical output:

# TYPE runtime_vul_severity_count gauge
runtime_vul_severity_count{image="nginx:1.21",severity="CRITICAL"} 1
runtime_vul_severity_count{image="nginx:1.21",severity="HIGH"} 2
runtime_vul_severity_count{image="nginx:1.21",severity="MEDIUM"} 0
runtime_vul_severity_count{image="nginx:1.21",severity="LOW"} 0
runtime_vul_severity_count{image="nginx:1.21",severity="UNKNOWN"} 0
# TYPE runtime_vul_last_scan_timestamp_seconds gauge
runtime_vul_last_scan_timestamp_seconds 1749081600
# TYPE runtime_vul_images_scanned gauge
runtime_vul_images_scanned 3
# TYPE runtime_vul_scan_errors_total counter
runtime_vul_scan_errors_total 0
Metric Type Labels Meaning
runtime_vul_severity_count gauge image, severity Vulnerabilities found per image and severity in the most recent scan. The full per-image set is replaced atomically each cycle, so images no longer running drop out — no stale series.
runtime_vul_last_scan_timestamp_seconds gauge – Unix timestamp of the last completed scan cycle. Alert on staleness with time() - runtime_vul_last_scan_timestamp_seconds.
runtime_vul_images_scanned gauge – Number of images scanned in the last cycle.
runtime_vul_scan_errors_total counter – Monotonic count of per-image scan errors since startup.

The exposition is hand-rendered over the standard-library HTTP server — no client library is pulled into either image, keeping the distroless Go binary and the hardened Alpine Python image dependency-free.

Securing the endpoint

Protection is layered and opt-in — compose the tier that fits your scrape path. All options are applied at startup (a restart is required to change them); a misconfiguration (e.g. an unreadable cert) is logged and the scanner keeps running with the endpoint disabled rather than crashing.

Tier Config Effect
Basic auth metrics_basic_auth_user + metrics_basic_auth_password /metrics returns 401 without valid HTTP Basic credentials (compared in constant time).
TLS metrics_tls_cert_file + metrics_tls_key_file Serves HTTPS. A self-signed cert works when the scraper sets insecure_skip_verify / tls_config.insecure_skip_verify: true ("tls verify=false"); a CA-signed cert lets the scraper fully validate the certificate.
Mutual TLS add metrics_tls_client_ca_file Requires and verifies a client certificate signed by the given CA (RequireAndVerifyClientCert) — the strongest tier.
# Example: TLS + Basic auth (certs mounted read-only under /config)
metrics_basic_auth_user: "prometheus"
metrics_basic_auth_password: "change-me"
metrics_tls_cert_file: "/config/tls/metrics-cert.pem"
metrics_tls_key_file: "/config/tls/metrics-key.pem"
# metrics_tls_client_ca_file: "/config/tls/client-ca.pem"   # add for mutual TLS

Cert file permissions: the container runs as UID 65532, so any mounted cert/key must be readable by that user (e.g. world-readable cert, and a key the runtime user can read). A key the daemon cannot read is reported as Failed to start metrics server and the endpoint stays down.

A matching Prometheus scrape config:

scrape_configs:
  - job_name: runtime-vul-detect
    scheme: https
    basic_auth: { username: prometheus, password: change-me }
    tls_config: { insecure_skip_verify: true }   # or `ca_file:` for full validation
    static_configs: [ { targets: ["127.0.0.1:9090"] } ]

./verify_metrics.sh [python|go|all] (repo root) is an integration check that brings a variant up with all three tiers enabled — generating a throwaway PKI — and asserts the endpoint enforces Basic auth, TLS, and mutual TLS. All generated material lives under <variant>/config/tls/ (gitignored) and is removed on exit.

Technical evaluations

Distroless base image

Google's gcr.io/distroless/python3-debian12 is not a viable drop-in replacement for python:3.13-alpine for this project. Key blockers:

Issue Detail
Python version Distroless is pinned to Python 3.11 (Debian 12 package). No official 3.12/3.13 variant exists (issue #1703).
ctypes missing The base python3-debian12 image does not include ctypes, which Cython-compiled .so files may depend on. The python3-plus variant adds it but also adds a shell.
musl vs glibc Our current Alpine base uses musl libc. Cython .so files compiled against musl cannot run in a glibc (Debian) environment — requiring a full rebuild.

Conclusion: Distroless Python would require switching the entire build chain from Alpine/musl to Debian/glibc, pinning to Python 3.11, and validating Cython compatibility. The security benefit (no shell, no package manager) is partially achieved today by stripping .py source via Cython and using Alpine's minimal attack surface. Revisit when distroless officially supports Python 3.13.

Go migration

Status update: a Go implementation now ships in go/ (see Implementations above). It shells out to the Trivy binary — the lower-risk approach described below — and is published as 1.0.0-go-distroless. The notes below remain the rationale and the outlook on deeper Trivy integration.

Rewriting the orchestrator in Go is meaningful long-term but not urgent today.

Concrete benefits of Go:

Dimension Python (current) Go (future)
Container image ~250 MB ~30–50 MB (single static binary)
Memory at idle ~60–80 MB ~8–15 MB
Startup time ~500 ms ~20 ms
Cross-compilation Requires separate environments GOOS=linux GOARCH=arm64 go build
Runtime dependency Python interpreter required None (statically linked)

Trivy Go library: Trivy exposes internal packages (pkg/commands/artifact, pkg/fanal/artifact/image) that can be called directly without a subprocess, eliminating the per-scan process fork and DB reload overhead. However, Trivy's internal API is not declared stable — upgrades require code changes. A safer approach is to use Trivy's --server mode (gRPC), which provides a stable interface while keeping Trivy as a separate process.

Realistic effort: A minimal Go rewrite that shells out to the Trivy binary (equivalent to today's Python behaviour) takes 1–2 weeks. Embedding Trivy as a Go library with proper DB caching and gRPC server integration is 4–8 weeks. The subprocess approach is lower risk and keeps pace with Trivy's CLI contract.

Recommendation: Migrate when adding Kubernetes support — Go's client-go and the Kubernetes informer pattern are significantly better suited to cluster-level workload watching than Python's kubernetes client.

Roadmap

Feature Status
Core cron scanning with JSON log output Done
docker ps -a image discovery (running + stopped) Done
Multi-file ignore with startup duplicate detection Done
Config-driven timezone (IANA) Done
Cython-compiled production image (no .py source at runtime) Done
Unit test suite with injected logger (no patching) Done
Go implementation — distroless image, stdlib Docker client (no SDK) Done
Docker socket proxy — GET-only allowlist, internal bridge network, no direct socket mount Done
Prometheus /metrics endpoint — severity counts as gauges Done
Webhook / alerting on new CRITICAL findings Done (opt-in)
Private registry support (TRIVY_USERNAME / TRIVY_PASSWORD) Planned
Hot-reload config on file change without container restart Done
Event-driven scanning on image pull (subscribe to Docker /events API) Done (opt-in)
Docker Compose project / service labels in scan_result log lines Planned
Scan state persistence across container restarts (eliminates webhook false positives on restart) Planned
Multi-arch image builds (linux/amd64 + linux/arm64) Planned
Kubernetes support (watch Pods, scan images via k8s API) Planned
Go rewrite with Trivy server mode (lower per-scan overhead) Evaluating

Webhook alerting on new findings

By default the scanner is detection-only — every cycle emits the full finding set as JSON, and acting on a result (paging, ticketing, Slack) is delegated to the downstream pipeline (Alloy → Loki alert rules, or a SIEM). On top of that, an opt-in webhook fires native alerts on newly appeared CRITICAL/HIGH findings: the scanner keeps an in-memory set of the CVE IDs seen for each image in the previous cycle, diffs the current findings against it, and POSTs only the CVEs that are new since the last scan — so a freshly disclosed critical in a long-running image triggers a notification without requiring a log backend to be wired up first.

The webhook is disabled by default (webhook_enabled: false) → behaviour is exactly today's, with zero HTTP. When enabled, results are diffed against the existing severity filter (no separate severity key); the first time an image is seen, all of its findings count as new (alert-on-everything on first scan). Delivery is best-effort and never breaks the scan loop: the configured webhook_timeout_seconds is applied per request, and any error (timeout, connection refused, non-2xx) is logged as a structured webhook_error event and the cycle continues. HTTP uses the standard library only (no third-party dependencies). The webhook_auth_token is never logged. webhook_* settings are applied at startup; changing them requires a container restart.

Set webhook_enabled: true and webhook_url in config.yaml, optionally with webhook_auth_header + webhook_auth_token (e.g. Authorization / Bearer …). One POST is sent per affected image per cycle, only when that image has new findings.

POST <webhook_url>
Content-Type: application/json
<webhook_auth_header>: <webhook_auth_token>   # only when both are configured
{
  "timestamp": "2026-06-06T12:00:00Z",
  "event": "new_vulnerabilities",
  "image": "python:3.8-slim",
  "image_id": "sha256:abc123...",
  "containers": ["a1b2c3d4e5f6"],
  "new_vulnerabilities": [
    {
      "id": "CVE-2021-1234",
      "package": "openssl",
      "version": "1.1.1k-r0",
      "fixed_version": "1.1.1l-r0",
      "severity": "CRITICAL",
      "title": "Memory corruption in openssl",
      "target": "python:3.8-slim (debian 11.7)"
    }
  ],
  "new_severity_counts": { "CRITICAL": 1, "HIGH": 0 }
}

The per-vulnerability object shape is identical to the vulnerabilities[] entries in the scan_result log line. new_severity_counts counts only the new vulnerabilities and always carries the CRITICAL and HIGH keys (a bucket may be 0). The Python and Go variants produce byte-for-byte identical config keys and POST bodies.

./verify_webhook.sh [python|go|all] (repo root) is an integration check that brings a variant up against the deliberately vulnerable test target with the webhook enabled — spinning up a throwaway receiver — and asserts a new_vulnerabilities POST with at least one CRITICAL/HIGH is delivered, then re-runs with webhook_enabled: false and asserts no POST arrives. All generated material lives under <variant>/config/webhook/ (gitignored) and is removed on exit.

Event-driven scanning

The cron schedule answers "is anything I'm already running newly vulnerable?" — but a freshly pulled image or a container started between ticks waits for the next cycle before it is ever scanned. Event-driven scanning closes that window: an opt-in listener subscribes to the Docker /events stream and reacts to two event kinds — an image pull (type=image, action=pull) and a container start (type=container, action=start). For a container start, the image ref is taken straight from the event payload (Actor.Attributes.image); no extra Docker API call is made.

Each qualifying event triggers a targeted scan of only that image — not a full re-scan of every container. The targeted scan reuses the regular per-image path, so it emits the same scan_result log line, feeds the webhook new-finding diff, and updates that image's per-image severity gauges. It deliberately does not touch the cycle-level bookkeeping: the health marker, runtime_vul_images_scanned, and runtime_vul_last_scan_timestamp_seconds belong to the cron cycle. Event-driven scanning therefore complements the cron schedule — it does not replace it.

Bursts are debounced: image refs observed while events keep arriving are collected into a set, and one scan per distinct ref fires only once the stream has been idle for events_debounce_seconds (default 5). A docker compose up that starts ten containers from three images results in three targeted scans, not ten. The same exclusions as the cron path apply — ignore.yaml image patterns are honoured and the scanner never scans its own image. Trivy executions are serialised internally, so an event scan and a concurrent cron cycle never run Trivy against the shared cache at the same time. If the /events stream drops (daemon restart, proxy hiccup) the listener logs a structured event_stream_error and reconnects with a small backoff.

The listener is disabled by default (events_enabled: false) → no /events connection is opened at all, exactly today's pure-cron behaviour. Enable it in config.yaml:

events_enabled: true
events_debounce_seconds: 5   # idle window before coalesced refs are scanned

events_* settings are applied at startup; changing them requires a container restart. Because the scanner reaches Docker only through the GET-only socket proxy, the proxy's -allowGET allowlist must also permit the events endpoint — each variant's shipped docker-compose.yml includes the events entry (both the API-versioned /v1.NN/events form used by the Python SDK and the bare /events path used by the Go client); without it the stream is blocked and the listener will log reconnect errors.

References

About

Cron-scheduled Docker host vulnerability scanner using Trivy — structured JSON logs for OpenTelemetry pipelines.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages