A catalog service is a reusable base definition for an application. A configured
service uses a catalog service and overrides only what differs. A service file
lives under paths.services, which is what marks it as a configured service — no
kind: is needed (see
configuration: a document's kind is derived from its
location).
name: apache-main
uses: apache
variables:
health_path: /health
watches:
restart-if-http-failed:
check:
url: "http://${host}:${port}${health_path}"The packaged catalog covers common service families such as web servers,
databases, container runtimes, NFS/libvirt helpers, observability daemons such
as Rsyslog, and hardware/system services.
In the source tree this is catalog/; in packaged builds Sermo reads the catalog
directory compiled into the binary. Catalog profiles define variables,
preflight, processes, watches, stop_policy, remediation policy and rules so a
configured service usually only sets a few overrides. High-impact catalog
services such as databases, caches and queues may carry stricter local policy
settings than the global defaults, with longer cooldowns, rate limits and
backoff to avoid restart loops.
The packaged snmpd profile runs its unauthenticated local SNMP protocol
probe only when the first rocommunity line in /etc/snmp/snmpd.conf grants
the community to localhost (no source, or a default/loopback one). An agent
that restricts v1/v2c to remote manager networks, or serves SNMPv3 users only,
cannot answer that probe, so it keeps just the init-service check; add an
authenticated site-specific SNMP check when protocol availability itself is
required there.
- Categories
- Library services
- Reload on config change (reload_on_change)
- App dependencies (apps)
- Metadata fields
- Built-in variables
- OS-specific blocks (os:)
- control: libvirt — QEMU/libvirt virtual machines
- control: docker — Docker containers
- Verified restart
- also_service — auxiliary init units
- also_apply — cascade to other services
- processes: by executable or cmdline
- Stopped-state invariants (stop_policy)
- Unclaimed control-group members (reap)
- pidfile: and pidfiles: shorthand (selectors + health checks)
- socket: shorthand (gated health check)
- lockfile: shorthand (gated health check)
- Versioned services
- Service unit
- Cloning
- Multiple instances of one application
- Catalog health policies
- Disabling and deleting inherited entries
- Monitoring flag
- Blocking operations while clients are connected
- Long-running database statements
- PostgreSQL replication watches
- Exim hints database maintenance
- Exim mail-volume alerts
- Grafana Alloy saturation restarts
- File-descriptor alerts (restart-if-fds-high)
- Auxiliary commands
Catalog documents are grouped by the subdirectory they live in under the packaged catalog root:
catalog/
services/ # long-running services (apache, nginx, mariadb, ...)
apps/ # installed tools/runtimes (java, perl, sqlite, go, git, ...)
libs/ # shared libraries used as restart triggers (glibc, pam)
patterns/ # output-analysis rule sets referenced by a check's analyze: block
The directory sets the catalog category (service / app / library /
patterns) and therefore the document's kind (service / app / lib /
patterns), so a top-level kind: is redundant and omitted; files placed
directly in the packaged catalog root are rejected. Use one YAML file per catalog
document: one service, app, lib or pattern in each file.
sermoctl services catalog, sermoctl apps and sermoctl libs list each
category, showing which are installed, the version their version command
reports, and whether they resolve without error (add all to include the
not-installed). Plain sermoctl services lists the configured service instances
(under paths.services) instead — see cli.md.
sermoctl patterns catalog lists every pattern set and its rule count, and
sermoctl patterns only the sets configured services use (see the analyze:
block in rules.md).
Catalog documents may declare aliases: [...] for distro or package names that
operators naturally type. For example, the canonical catalog service
name: apache can carry aliases such as apache2 and httpd, so a configured
service may write uses: apache2 while resolving to the same catalog profile. A
configured service may also declare aliases; sermoctl normalizes those aliases to
the canonical configured service name before status, start, stop, restart,
reload, monitor, SLA and process/lock commands. Catalog aliases are also usable
as service names only in the conservative one-service case where a configured
service has the same name as the catalog service, such as name: smb,
uses: smb, with catalog alias samba.
The gluster-ta-volume profile requires a non-empty volume file at
/var/lib/glusterd/thin-arbiter/thin-arbiter.vol. Override variables.config
when the systemd unit uses another path. Installing the GlusterFS binary alone
does not configure the arbiter; Sermo blocks operations that require preflight
when this file is missing or empty. sermoctl repair does not generate it.
The packaged arbiter unit and glusterd both use port 24007. Check the listening
addresses and the clients' arbiter endpoint before configuring them on one host.
Changing only the arbiter's listening port does not update its clients. The
profile monitors the arbiter unit; an active unit does not establish volume
quorum or replication health. Follow the
GlusterFS thin-arbiter setup
for the storage configuration.
A library service describes a shared library so configured services can restart when it is upgraded. It only needs identity plus the file to watch:
name: glibc
display_name: "GNU C Library"
description: "Standard C library (libc)"
variables:
binary: "/lib64/libc.so.6" # the file watched for changes (and its version)
preflight:
file: { type: file, path: "${binary}" }Set a top-level interval on an app or library profile to override the global
engine.artifact_interval (default 5m) used for artifact inspection.
When a service subscribes through restart_on_change.libraries, Sermo also
adds that library file as a required preflight check for start, restart, reload
and resume; a missing, non-regular or empty library file blocks the operation.
A configured service (or catalog service definition) opts in with
restart_on_change. Packaged catalog services that link versioned apps declare
the app form by default; custom services can use the same shape. paths is for
configuration files that require a full restart rather than a reload:
restart_on_change:
config: true
version: true
paths:
- ${config}
libraries: [glibc, pam]
apps:
containerd:
level: minor
messages:
path: "${display_name} will restart after config change: ${change.path}"
app: "${display_name} will restart after version change of ${change.app}: ${change.old_version} -> ${change.new_version}"On resolution this desugars into one remediation rule per path that restarts the service when that config path changes, one rule per library that restarts the service when that library's file changes, and one rule per app that restarts the service when the linked app's version changes at the selected level. Each generated rule alerts first, inheriting the normal rule/global notifiers, then runs the restart action through the safe operation engine:
rules:
restart-on-change-config-1:
type: remediation
if: { changed: { path: /etc/containerd/config.toml } }
then:
actions:
- type: alert
message: "containerd will restart after config change: ${change.path}"
- type: restart
restart-on-change-glibc:
type: remediation
if: { changed: { library: glibc, path: /lib64/libc.so.6 } }
then:
actions:
- type: alert
message: "containerd will restart after library change: ${change.library} (${change.path})"
- type: restart
restart-on-change-containerd-version:
type: remediation
if: { changed: { app: containerd, level: minor } }
then:
actions:
- type: alert
message: "containerd will restart after version change of ${change.app}: ${change.old_version} -> ${change.new_version}"
- type: restartThe daemon samples these paths and linked app versions at engine.artifact_interval
(or the applicable local interval). A service may evaluate rules more often,
but it reuses that sample; detection is therefore delayed by at most one artifact
cadence plus one scheduler tick.
The optional config and version booleans are inherited permissions. When
absent they default to allowed, preserving the current service behavior.
config: false suppresses generated paths restart rules. version: false
suppresses generated apps and libraries restart rules. Global defaults may
set only these two booleans:
defaults:
restart_on_change:
config: false
version: trueA catalog service or configured service may override either flag in its local
restart_on_change block.
Three mechanisms restart or reload a service when something it depends on changes. Each has a permission gate, and every gate is settable per host:
| Mechanism | Trigger | Gate | Notifies |
|---|---|---|---|
restart_on_change |
app version, library, config path | config: / version: |
yes — alert, then restart (app version and library changes are graded warning) |
reload_on_change |
config path | config: |
no — reload only, no alert action |
restart_on_stale_binary |
binary replaced on disk | the flag itself | yes — alert, then restart (graded warning) |
restart_on_fds_high |
a process above fds_limit of its open-files limit for 3 minutes |
explicit true (default false) |
yes — alert; restart only when enabled |
Two levels of granularity, both on the host, neither requiring a catalog edit:
# /etc/sermo/sermo.yml — the whole host
defaults:
restart_on_change: { version: false } # no version-driven restarts here
reload_on_change: { config: false } # no config-driven reloads either
restart_on_stale_binary: false# /etc/sermo/services/nginx.yml — just this service, on this host
name: nginx
uses: nginx
restart_on_change: { version: true } # …except nginxThe merge is deep, which is what makes the host level usable: setting only
the gate in defaults: folds into the catalog's block instead of replacing it,
so the paths, apps and messages the catalog ships survive. Precedence is
defaults: < catalog < the host's per-service file — see
Resolution order. A scalar the catalog sets
explicitly (as the docker and OVS profiles do for restart_on_stale_binary)
therefore beats a host-wide default; override it in the per-service file
instead.
Global defaults: accepts only the gates, never paths/apps/libraries/
messages: a host decides whether these restarts happen, the catalog decides
what triggers them.
messages is optional and local to the service or catalog service. It accepts
path, app and library templates. The templates are expanded like normal
service strings first (${display_name}, ${config}, …), then rule runtime
placeholders such as ${change.path}, ${change.app}, ${change.old_version}
and ${change.new_version} are filled when the alert is emitted.
The restart runs through the normal safe engine (guards, cooldown, max_actions),
and the change is acknowledged once the restart succeeds, so it fires once per
upgrade rather than every cycle. Referenced library names must be library
services. Referenced app names must also appear in the service's apps: list,
and the app must provide a version or version_short command. App levels are
major, minor and patch (default for the short form apps: [containerd]).
If the app binary or version command is broken, Sermo treats the version sample
as invalid, does not update the version baseline, and does not restart the
service.
Many services re-read their configuration without a restart — systemd
(systemctl daemon-reload), nginx (nginx -s reload), named (rndc reload),
rsyslog, … reload_on_change watches config files/directories and, when one
changes, runs the reload action instead of a disruptive restart:
# catalog/services/systemd.yml
reload:
command: ["systemctl", "daemon-reload"]
when: always
reload_on_change:
paths: [/etc/systemd/system, /lib/systemd/system]On resolution this desugars into one remediation rule per path:
rules:
reload-on-change-1:
type: remediation
if: { changed: { path: /etc/systemd/system } }
then: { action: reload }Note the single action. Unlike restart_on_change, the generated rule carries
no alert action, so a reload sends no notification — it is recorded as an
operation-result event and nothing more. That is deliberate for a non-disruptive
reload; if you want to be told, add your own alert rule on the same changed:
condition. Set reload_on_change: { config: false } to suppress the generated
rules entirely, per service or per host.
The reload action runs through the same safe engine as restart but in
place: it runs preflight first (so an invalid config — caught by the
service's config check — blocks the reload), reloads, then verifies health.
reload is also a valid rule action on its own (then: { action: reload }) and
is blocked by guards that list reload, like any other service action.
What "reload" runs. By default it is the backend per-unit reload —
systemctl reload <unit> (which runs the unit's ExecReload, e.g. nginx -s reload) or OpenRC's init-script reload. A catalog service can override this with
reload.command when the reload is not a per-unit operation — systemd
itself reloads with systemctl daemon-reload, not systemctl reload systemd:
reload:
command: ["systemctl", "daemon-reload"]
when: alwaysIf the init backend reports no reload support and the service has no valid
reload.command or reload.signal fallback, Sermo rejects the reload action
before execution. The CLI reports the unsupported reload and the web UI disables
the reload button through can_reload=false.
Some services reload in place (e.g. sshd, snmpd, proftpd, prometheus,
loki re-read their config on SIGHUP) but their systemd unit defines
no ExecReload, so systemctl reload <unit> fails — even though the service
itself supports it (the same service under OpenRC usually does reload, via an
init-script reload() that sends the signal). The reload: block closes that
gap: it declares a native reload Sermo performs itself, by signalling the
service's main process or running a command.
reload:
signal: HUP # send this signal to the main process (HUP, USR1, USR2, …)
when: auto # auto (default): use the init's reload if the unit/script
# has one, otherwise do this; always: never use the init,
# always do this
# or, instead of a signal, a command:
reload:
command: ["nginx", "-s", "reload"]
when: autowhen: auto(default) asks the backend whether it can reload — systemd'sCanReload(the unit has anExecReload), or an OpenRC init script that definesreload. If it can, the init reload runs; if it can't, Sermo runs the native reload. So the same catalog service definition reloads correctly on a host whose unit exposes reload and on one whose unit doesn't.when: alwaysalways runs the native reload and never the init's — the right choice for reloads that are not per-unit operations. A barereload: { command: [...] }defaults towhen: auto, so setwhen: alwayswhen the command must always run.- Signal target. The signal goes to systemd's
MainPID, or — on OpenRC, or any unit with no MainPID — to the PID in the service'spidfile:. The pidfile fallback is only used when that PID also matches aprocesses:selector with exactexeanduser; a stale pidfile must not signal an unrelated process. A signal reload with neither target available fails. Services without pidfile metadata reload by signal only on systemd; on OpenRC they rely on the init script's ownreload(when: auto).
Before shipping or changing a catalog service with reload.signal, verify every
init backend listed in service: and every fallback Sermo may use. Do not check
only the platform where the profile was first written.
- Inspect the real packaged init definitions. For OpenRC, read
/etc/init.d/<unit>and the matching/etc/conf.d/<unit>; for systemd, read the unit and its reported reload/PID metadata. - Record whether the init backend can reload by itself. With
when: auto, Sermo prefers the backend reload when systemd reportsCanReload=yesor the OpenRC script definesreload(). If a host lacks that path, Sermo's native fallback must still be safe. - For any OpenRC-capable
reload.signal, declare a canonical/run/...pidfile:candidate and aprocesses:selector with exactexeanduser. The executable must be the resolved/proc/<pid>/exepath (usually through the linked app's binary variable), and the user should be a service variable so local packaging differences can override it. - If OpenRC scripts differ by distribution, encode the real pidfile candidates
as a list or an
os:branch. Do not ship a single path that was verified on only one distro. - If a backend has no pidfile or no trustworthy
exeplususeridentity, do not rely onreload.signalfor that backend. Use an argvreload.command, or rely only on the init backend's reload when every configured backend validates. - Run the catalog validation tests for both init backends before release.
Useful host checks:
sermoctl backend
systemctl cat <unit>
systemctl show -p CanReload -p MainPID -p PIDFile -p User <unit>
sed -n '/^reload()/,/^}/p' /etc/init.d/<unit>
grep -E '^(command|command_user|pidfile|.*PIDFILE)=' /etc/init.d/<unit> /etc/conf.d/<unit>
readlink -f /usr/sbin/<service>
namei -l /run/<service>.pidUseful catalog audit while developing:
go test ./internal/config -run 'TestRealCatalog(AllServicesValidate|ReloadServicesResolve)$' -count=1The reload path chosen by the backend or by reload: is what the reload
action, reload_on_change, the sermoctl reload <svc> command and the web UI
reload button all run. It is a service-control concept: it applies to services,
not to host watches, which observe host metrics and fire hooks rather than
reload a unit.
A service can link one or more apps from catalog/apps (java, openssl,
perl, …). An app owns the tool's binary, health and version checks.
Link them with apps::
# catalog/services/tomcat.yml — Tomcat runs on the JVM
apps: [java, "tomcat-${version}"]On resolution each linked app's preflight checks are injected into the service's
preflight under keys namespaced by the app name (<app>-<check>), carrying the
app's own variables.binary path, health probe and version command. Link an app
only when the service action itself requires it. For example, Backrest can be
monitored and restarted without restic; restic is required by a backup
operation, which reports its own error if the binary is absent. Likewise,
Samba's winbindd belongs in an enable_if-guarded process/watch, not in
apps, because it depends on the host's Samba configuration.
When a service links several required apps, each one's checks stay distinct:
preflight:
java-binary: { type: binary, path: /usr/bin/java }
java-health: { type: command, command: ["/usr/bin/java", "-help"] }
java-version: { type: command, command: ["/usr/bin/java", "-version"] }App variables are also available to the service. They are always exposed with a
normalized app-name prefix (${java_binary}, ${php_fpm_binary}, ...). If the
service links exactly one app, those variables are additionally available without
the prefix as defaults, so service-specific checks can use ${binary} while the
app keeps ownership of the actual path. Local variables: entries on the catalog
service or configured service override either form; when several apps are linked,
use the prefixed names.
Because they run in preflight, a missing or wrong-version runtime fails the
service's preflight, which blocks start/restart/reload/resume (a
preflight-failed operation never executes the action) — you do not start,
restart, reload or resume a service whose runtime is absent.
The link is many-to-many: a service lists several apps, and one app is shared by
every service that lists it. Validation reports an apps: entry that does not
resolve to a catalog app, so dangling runtime links are caught before deployment.
The service keeps its own variables.binary,
version and config checks (the config test is always service-specific,
never moved to an app). Referenced names must be app services.
A catalog service or configured service may carry optional human-facing metadata:
name: mariadb
display_name: "MariaDB" # pretty label; falls back to name when absent
description: "..." # free-text note; shown verbatim, nothing when absent
category: "database" # optional WebUI grouping/filter label
type: "database" # optional free-form classification; recorded, not acted onThese fields are optional and behave differently when missing:
display_nameis the label used wherever Sermo shows the catalog entry to a human (e.g.sermoctl services catalog,sermoctl appsand the Web UI). When it is absent or blank, Sermo falls back toname. Set it only when it adds something overname— a proper brand (MariaDB,PostgreSQL,OpenSSH) or a version (PHP-FPM 8.3). If the display name would just repeatname, leave it out and let the fallback apply.descriptionis an optional free-text note. It has no fallback: when it is absent, nothing is shown for it — Sermo never substitutesname. Use it for a real sentence, not a restatement of the name.categorygroups and filters Services, Installed applications and Installed libraries in the WebUI. When absent or blank, services useservice, apps useappand libraries uselibrary.typeis an optional free-form classification label (e.g.database,cache,queue,webserver,appserver,tunnel) used in the catalog to organize entries. It is recorded but not currently consumed by the engine and has no effect on monitoring, grouping or remediation.
display_name, description and category must be strings if present;
validation rejects non-string values.
The variables in the table below are always available during resolution
without being declared under variables — so a catalog service can parameterize
human-facing strings (and paths) instead of hardcoding them:
rules:
block-restart-during-maintenance:
type: guard
blocks: [restart, stop]
then:
action: block
message: "${display_name} maintenance is active" # → "MariaDB maintenance is active"
variables:
binary: "/usr/bin/qemu-system-${arch}" # → /usr/bin/qemu-system-x86_64
preflight:
binary: { type: binary, path: "${binary}" }An explicit variables entry of the same name always takes precedence over a
built-in. ${arch}/${os} are baked on load (everywhere — variable values
and app discovery paths included); the rest resolve per service, and
the runtime ones (${date}/${event}/${action}) only in rule message:
strings. The SERMO_ARCH / SERMO_OS / SERMO_HOST / SERMO_HOSTNAME /
SERMO_INIT / SERMO_USER environment variables override the matching built-in
(handy for testing or building config off-host).
${user} is a config-load built-in. It uses SERMO_USER when set, otherwise
the user running Sermo. It is intentionally separate from the runtime
engine.user_lookup resolver used for process selectors and kill_only_if; set
SERMO_USER when you need ${user} to be deterministic while generating or
validating config off-host.
| Variable | Value | Resolved |
|---|---|---|
${name} |
the resolved service name | resolution |
${display_name} |
the display name (falls back to name) | resolution |
${service} |
the primary unit name for the active backend | resolution |
${host} |
hostname (SERMO_HOST override) |
resolution¹ |
${hostname} |
short hostname (SERMO_HOSTNAME) |
resolution⁵ |
${init} |
detected init system (SERMO_INIT) |
resolution |
${user} |
Sermo's user (SERMO_USER override) |
resolution⁴ |
${pidfile} |
conventional /run/<unit>.pid |
resolution⁴ |
${port} |
the top-level port: field (when set) |
resolution³ |
${arch} |
machine architecture (SERMO_ARCH) |
load (baked) |
${os} |
os-release id (SERMO_OS) |
load (baked) |
${date} |
event timestamp (RFC3339) | runtime² |
${event} |
the firing rule's name | runtime² |
${action} |
the action taken (restart/start/stop/reload/resume) | runtime² |
¹ ${host} only applies when the service does not define a host variable (a
bind address like 127.0.0.1); an explicit host always wins.
⁵ ${hostname} is the short hostname — the first label before the first dot
(node1 on node1.example.com) — distinct from ${host} (which keeps the full
detected hostname / bind-address fallback). Use it for systemd instance units
keyed by host identity, e.g. service: "ceph-mon@${hostname}" → ceph-mon@node1.
For numeric multi-instance services (e.g. one OSD per device) use a %n service
template whose service: carries ${n}. Sermo materializes ceph-osd0…N from
active units such as ceph-osd@0.service, then links the generic ceph-osd app
for binary validation. An explicit hostname variable (or SERMO_HOSTNAME)
wins.
⁴ ${user} and ${pidfile} are fallbacks: a service's own user (a service
account such as www-data) or pidfile variable always wins. Put the pidfile
variable in the service-level pidfile: "${pidfile}", and use user: "${user}"
inside any processes: selector that should be tied to the service account.
Runtime paths in Sermo config use the canonical /run spelling. Do not write
new /var/run pidfiles, sockets or lockfiles in catalog services, generated
services or examples. Linux keeps /var/run as compatibility for /run, and
older init scripts, service managers or packaged configs may still report that
spelling; detected paths should be normalized to /run/... before they are
committed to config.
Before adding a new runtime path, check whether it or a parent directory is a
symlink (readlink -f <path> or namei -l <path>), then record the canonical
target path rather than the alias.
² ${date}/${event}/${action} are substituted when the worker emits a rule
message, so they belong in message: strings — e.g.
message: "[${host}] ${service}: ${event} → ${action} at ${date}". Elsewhere they
stay literal.
³ ${port} mirrors a top-level port: field on the configured service (or catalog
service), so an instance can set its listen port once and have every ${port}
reference resolve to it:
name: db-inst2
uses: dbserver
port: 3307 # → ${port} everywhere in the catalog serviceUnlike the other built-ins it has no fallback: declare port: (or a
variables.port, which wins) wherever ${port} is used, or resolution reports
${port} as undefined. This is the first-class equivalent of putting port
under variables: (as the multi-instance example below still shows).
Beyond the ${os} string, an os: key anywhere in a document selects a whole
sub-block by OS. The block for the detected OS (or a default block) is merged
into its parent and the rest discarded — at load, before resolution. It is not
limited to the service block; use it in checks, processes, policy, variables, anywhere:
service:
os:
gentoo: { systemd: [apache], openrc: [apache] }
debian: { systemd: [apache2], openrc: [apache2] }
watches:
http:
check:
type: http
timeout: 5s # kept for every OS
os:
gentoo: { url: "http://localhost/gentoo-health" }
debian: { url: "http://localhost/debian-health" }
policy:
os:
debian: { cooldown: 1m }
default: { cooldown: 9m } # used when the OS has no branchSiblings of os: are preserved and the selected branch merges over them. os is
reserved as a selector key wherever its value is a map.
A branch may also be a list or scalar instead of a map. When os: is the only
key in its parent, the selected branch replaces the value (rather than merging),
which is handy for OS-specific candidate lists such as pidfile paths:
pidfile: # the resolved value becomes the OS's list
os:
fedora: [/run/postgres.pid]
gentoo: [/run/postgres${port}.pid, /run/postgres.pid]
default: [/run/postgres.pid]The service-level pidfile: accepts a single path or a list of candidates.
Discovery tries them in order and uses the first that points at a running
process, so per-OS or versioned pidfile locations all resolve without personal
config. Use pidfiles: instead when one service intentionally owns several
resident processes that each have their own pidfile.
The rest-server profile discovers its resident process using the linked
application's rest_server_binary and an exact user (default root), including
on OpenRC installations without a pidfile. Set variables.user to the account
that runs your instance; set variables.rest_server_binary when it lives outside
the discovered binary directories. Both values must match for process and FD
metrics to be attributed to the service.
For oneshot services that do not keep a resident process (for example firewall
loaders or SNTP), set processes: {} explicitly. That prevents Sermo from deriving a
process selector from init metadata and keeps the WebUI from showing CPU/memory
process totals for a service that cannot have them.
This also omits the generated file-descriptor and stale-binary checks and rules. An
expect: active service check is appropriate only if the unit remains active
after completing. A systemd oneshot skipped by Condition* can be inactive
without failing; monitor its result with a read-only host watch instead:
name: boot-task-result
interval: 5m
check:
type: command
command: [systemctl, show, boot-task.service, "--property=LoadState,Result"]
expect_stdout:
op: "=~"
value: '^(LoadState=loaded\nResult=success|Result=success\nLoadState=loaded)$'
timeout: 5sThis checks the last result, not whether the task has ever run or how recently
it completed. LoadState=loaded also ensures a removed unit fails the check,
even if systemd still reports its default Result=success.
When a timer schedules the oneshot, make the timer the controlled unit: it
is what an operator enables, starts and stops, and it stays active between
runs, so the service check and start verification apply to it. Read the
oneshot's last result with the same command check as a service watch, and
watch the freshness of whatever the task writes to prove it keeps running. The
logrotate profile is the reference: logrotate.timer is the service, a
last-run watch reads logrotate.service's result, a monitor-only file watch
alerts when the state file is older than two scheduled runs, and the config
preflight runs logrotate --debug so a broken file under /etc/logrotate.d is
reported before the night it is skipped. OpenRC installations run logrotate from
cron.daily and have no unit, so the profile declares none for that backend.
A service can be controlled as a libvirt/QEMU virtual machine instead of a systemd/OpenRC unit:
name: vm-web01
control:
type: libvirt
uri: qemu:///system
domain: web01
socket: /run/libvirt/libvirt-sock # or /run/libvirt/virtqemud-sock on modular libvirt
watches:
vm:
check:
type: libvirt
socket: /run/libvirt/libvirt-sock
query: qemu:///system
params: { domain: web01 }
processes:
qemu:
exe: /usr/bin/qemu-system-x86_64
cmd: "web01|2b3f3d26-bb45-4b25-b65a-1e3ef86fc1a4"
user: qemucontrol.domain is the libvirt domain Sermo operates. uri defaults to
qemu:///system; socket defaults to /run/libvirt/libvirt-sock unless host
is set for a remote libvirt TCP connection. Modular libvirt deployments often
expose QEMU domains through /run/libvirt/virtqemud-sock; set socket to that
path when the monolithic socket is absent. uuid is optional and, when set,
Sermo looks up the domain by UUID instead of name.
The safe operation engine is unchanged: locks, guards, preflight, postflight, operation timeouts and remediation policy still apply. The primitive actions are libvirt operations:
startcreates/boots the defined domain (DomainCreate).stoprequests a graceful guest shutdown (DomainShutdown); it does not destroy the VM.restartis still Sermo's safe stop+start flow.pausesuspends a running domain in place (DomainSuspend, likevirsh suspend): its vCPUs stop and its memory stays resident.resumeresumes a paused domain (DomainResume).reloadis unsupported for VM domains unless a future service-specific mechanism is added.
Libvirt status maps to Sermo status as follows: running/blocked → active,
paused/pmsuspended → paused, shutoff/nostate → inactive, crashed →
failed, and shutdown (a guest still shutting down) → unknown, so a stop or
restart keeps waiting until the domain is shut off. The CLI and web UI still expose backend status=paused; the aggregated
service state is failed while monitoring is active, or stopped when Sermo
monitoring is paused.
Process discovery is intentionally explicit in this first VM integration. If you
want process metrics or residual-process reporting for the QEMU process, add a
restrictive processes: selector as above: exact exe and user plus a cmd
regex that narrows the shared QEMU binary to the intended domain or UUID. The
cmdline selector narrows discovery; residual signaling is still
authorized only by stop_policy.kill_only_if.
sermoctl wizard vm can generate this service shape from domains
detected through the local libvirt socket. It probes both
/run/libvirt/libvirt-sock and /run/libvirt/virtqemud-sock and writes the
socket it actually used into the generated service and check.
A service can control one libvirt virtual network (virsh net-start /
net-destroy territory) instead of a domain:
name: libvirt-net-default
category: virtual-network
control:
type: libvirt-network
uri: network:///system
network: default
socket: /run/libvirt/virtnetworkd-sock
guard_socket: /run/libvirt/virtqemud-sock
processes:
dnsmasq-root:
exe: /usr/bin/dnsmasq
cmd: '--conf-file=/var/lib/libvirt/dnsmasq/default\.conf([[:space:]]|$)'
user: root
delegated: truecontrol.network is the libvirt network name. Network RPC runs over
socket/uri (defaults /run/libvirt/libvirt-sock and network:///system,
which monolithic libvirtd accepts too); the guest-attachment guard below needs
domain APIs, which on modular libvirt live on a different daemon, so it
dials guard_socket/guard_uri (defaults: the network socket and
qemu:///system). host/port select a TCP endpoint exactly like
control: libvirt, shared by both sessions.
The safe operation engine is unchanged: locks, guards, preflight, operation timeouts and remediation policy still apply. The primitive actions are libvirt network operations:
startstarts the defined network (NetworkCreate).stopdestroys the network (NetworkDestroy) — but hard-refuses while any live guest has an interface on the network (matched by source network name or by the network's bridge, and counting paused and crashed guests: their taps stay attached; only a shut-off guest is free). Destroying such a network cuts guest connectivity and the taps do not reattach on the next start. No configuration option relaxes this guard, and an unverifiable guest blocks the destroy rather than being skipped. The guard sees only the domains ofguard_uri: guests of another driver or of a session URI on the same bridge are not inspected.restartis still Sermo's safe stop+start flow, so it inherits the guard.reload,pauseandresumeare unsupported for virtual networks.
Network state maps active → active and inactive → stopped/failed
following the usual monitoring semantics.
Why manage networks at all: libvirt spawns one dnsmasq pair per NAT network,
and that pair deliberately survives daemon restarts (see the packaged
virtnetworkd profile, where it is delegated). After a dnsmasq package
upgrade those processes keep running the replaced binary and only a
network restart renews them — so the generated network service is the one
target whose stale-binary finding a restart genuinely fixes. Attribute the
pair with delegated: true selectors like the example above: Sermo observes
it, libvirt owns its lifecycle.
The fleet installer generates one such service per active network with a libvirt-owned IP (the shape that spawns dnsmasq); bridge-mode networks are recorded as skipped because a restart renews nothing on them, and a host with no domain-API socket is skipped too — without one the attachment guard could not verify guests, and an unverifiable destroy target must not exist.
A service can be controlled as one Docker container instead of a systemd/OpenRC unit:
name: web-container
control:
type: docker
container: web
socket: /run/docker.sock
watches:
docker:
check:
type: docker
socket: /run/docker.sock
container: web
on_change: true
expect:
container.status: { op: "==", value: running }
container.health: { op: "==", value: healthy }control.container is the Docker container name or id Sermo operates. With no
socket or host, control uses /run/docker.sock; set socket for another
local socket, or set host and optional port/tls for a TCP Docker API
endpoint. control.interface is not supported for control; interface-bound
egress remains available on Docker checks.
The safe operation engine is unchanged: locks, guards, preflight, postflight, operation timeouts and remediation policy still apply. The primitive actions are Docker Engine API operations:
startcalls the container start endpoint.stopcalls the container stop endpoint with no Docker-side kill escalation; the request and subsequent graceful observation sharestop_policy.graceful_timeout(10 seconds when omitted or zero). If that deadline expires, Sermo records the request error, discovers residuals and applies its stop policy. The overall operation timeout still bounds every phase; expiring or cancelling it prevents further escalation and start.restartis still Sermo's safe stop+start flow.pausefreezes every process of a running container (POST /pause).resumeunpauses a paused container.reloadis unsupported for Docker containers unless a future service-specific mechanism is added.
Docker status maps to Sermo status as follows: running -> active, paused ->
paused, created/exited -> inactive, restarting/dead/removing -> failed.
The CLI and web UI still expose backend status=paused; the aggregated service
state is failed while monitoring is active, or stopped when Sermo monitoring
is paused.
For process metrics and residual-process reporting, Sermo reads the container's
State.Pid from Docker inspect and discovers that process tree. You normally do
not need a processes: selector for a controlled container. Residual signaling
is still authorized only by stop_policy.kill_only_if.
After a timed-out stop request, a container without selectors is considered
stopped only when a fresh process snapshot proves the previous process
generations exited and Docker confirms the container is inactive. An empty PID
lookup or incomplete process snapshot is insufficient.
After process exit, Sermo allows up to four seconds for Docker to publish its
inactive state, within the same operation deadline. Inspection errors or a
container that remains active block the following start.
The same evidence ends the graceful wait early when the container has already
stopped; a successful stop need not consume the entire grace period.
Reloading Sermo while a container is stopped preserves process observation for its next start. Service events, including automatic recovery, refresh the web status cache so a completed repair does not leave the previous inactive status displayed until the periodic backend refresh.
sermoctl wizard docker can generate this service shape from containers
detected through the local Docker socket. The generated check requires
container.status == running and container.health != unhealthy, so an
unhealthy container remains a failure even without a new state transition.
Containers without a Docker health check and those still in the health check's
startup grace period are not rejected by the health predicate. Configure
container.health == healthy explicitly when readiness must also be required.
sermoctl restart SERVICE always composes stop and start inside one operation:
one lock, one timeout, required preflight, guards, process identity verification,
postflight and one auditable result. Sermo does not invoke the init backend's
restart command. The former restart_policy setting is rejected; remove it
from defaults and service overrides when upgrading.
Before start or restart, Sermo compares the init state with a fresh process observation. An active resident service with no processes can be reconciled only when absence is demonstrated by configured exact executable/user identities and complete process discovery. Missing identity, unreadable procfs, an unresolvable user or uncertain pidfile data cannot authorize this recovery. Services declared without a resident process retain their existing lifecycle semantics: active without a PID is valid for them. Zombies do not count as running processes and cannot restrict discovery to a dead process tree; monitoring can still report them independently.
- Init active, daemon absent: OpenRC clears its stale started marker with
rc-service SERVICE zap; systemd first runssystemctl stop UNIT, then clears failed bookkeeping withsystemctl reset-failed UNIT. Sermo verifies that init is inactive before proceeding. A restart uses its normal stop phase for this reconciliation, without a second stop or graceful wait.reset-failedcannot deactivate an active unit. - Init inactive/failed, daemon alive: Sermo handles survivors under the same
stop_policyas a stop, excluding delegated workloads. Any surviving orphan or discovery error blocks a following start. A successful cleanup reconciles init before proceeding.
The stop phase runs the backend stop, waits up to graceful_timeout for fresh
process observations to confirm exit, handles residuals, reconciles init and
verifies the stopped state. It proceeds immediately when the backend has already
stopped the processes; process-free services instead wait for inactive init state.
The signal escalation limits term_timeout and kill_timeout also end early
when fresh discovery confirms that no residual remains. These are maximum grace
periods, not fixed delays added to every restart. A nonzero command result does
not by itself prove that stopping failed: when Sermo demonstrates that the daemon
is gone and init can be reconciled, it continues and records the command error
as a warning. Otherwise it fails without starting a replacement. A failed reset
or an init state that remains inconsistent is an operation failure.
Stop-command, auxiliary-stop and stopped-artifact warnings stay in the result
and audit event even if a later phase fails or times out. An error during
residual rediscovery stops further signal escalation and blocks start.
The start phase verifies both the init state and the expected resident process,
then runs postflight. A command error can be recovered only with a trusted live
process and an active init state; the error remains visible in the result.
An executable replaced on disk can justify stopping the old daemon when its
previous path and real user match the configured identity. It cannot verify a
successful start: that requires the current executable. When main declares an
exact executable/user identity, a surviving worker or an unrelated deleted
executable cannot stand in for that main process.
Postflight must still pass. Reload and resume errors are not recovered merely
because the service remains active: that does not prove their requested effect.
Cancellation and timeout stop the flow; recovery never creates a new deadline.
Socket/D-Bus activation can replace the daemon while stop is settling. Sermo accepts this only for an active systemd unit with verified backend-attributed processes from a new generation (PID plus start time), with none of the old non-delegated generations remaining anywhere in the observed process table. Leaving the unit's cgroup or process tree does not prove that an old daemon exited. It then verifies health without issuing a redundant primary start. Auxiliary units stopped by this operation are started again before health verification. Finding the unchanged daemon active is not proof of a restart.
Delegated workloads, including container shims and Gluster workers, remain
excluded from residual cleanup. Auxiliary units declared by also_service
participate in the same stop/start ordering below.
A service can name auxiliary init units of its own (a .socket, .timer,
companion unit) that are started/stopped together with the primary, in the
same operation. A restart composes those two operations. It mirrors the
service: shape (per-init lists, resolved for the active backend):
service:
systemd: [docker]
openrc: [docker]
also_service:
systemd: [docker.socket]These are plain init units driven directly by the service manager (not separate
monitored services — that is also_apply). They are started before the
primary in declaration order (strict — a failure aborts the operation before
the primary starts). On systemd, auxiliary .socket, .timer and .path
units stop before the primary, preventing activation during its shutdown.
Other auxiliary units stop after the primary. Each stop group runs in
reverse declaration order; stop failures remain best-effort warnings in the
operation result. Cancellation stops the sequence. reload touches the primary only. The primary's
guards, locks and preflight wrap the whole operation. Listing the primary unit in
also_service is rejected.
Where also_service acts on init units of this service, also_apply acts on
other Sermo services: when this service is started/stopped/restarted (by a
remediation rule, sermoctl or the Web UI), the same action runs on each listed
service through its own guarded operation.
also_apply: [nginx, varnish]- Dependency-aware order: on
start/restartthe primary acts first, then the additionals (a dependent comes up after what it depends on); onstopthe additionals act first, then the primary. - Each target keeps its own guards/locks/preflight (it runs its real
operation). A target's remediation cooldown and paused/
unmonitorstate are not consulted —also_applyis an explicit relationship. When a remediation cascade starts or restarts a target successfully, the target also acknowledges its pendingchanged:baseline, as its own restart would, so its ownrestart_on_changerule does not restart it a second time. - Best-effort & loop-safe: a blocked target is retried once; if it remains
blocked, it is reported by a non-fatal
cascadeevent. A failed or unresolved target is also reported and makes the overall cascade result fail, while the remaining targets are still attempted. Cycles are cut by a visited set. These ordering, retry and result rules are identical for remediation, CLI and Web UI actions. - Entries must be configured services and must not include the service itself.
sermoctl start|stop|restart <svc> --no-cascadeacts on exactly one service. Without it, text output adds onecascade <target>: <action> <status>line per target; with--jsonthe single JSON result gains acascadearray of{service, action, status, message, error}objects (the final outcome of a retried target) instead.sermoctl reload <svc>,sermoctl pause <svc>andsermoctl resume <svc>act on the primary only (no cascade). Usesermoctl daemon reloadto reload the runningsermodconfiguration. In the web UI the per-service reload button is enabled only when the service isactiveand Sermo reportscan_reload=truefrom either the init backend (ExecReload/OpenRCreload) or a validreload:fallback; VM and container rows carry one pause toggle (⏸): it pauses anactivetarget and, latched while the target ispaused, resumes it. A paused row is tinted in its own colour and its state readspaused.
also_apply (other services) and also_service (this service's init units) are
complementary; a service may use both.
A processes: selector matches a process by the AND of the fields you set;
at least one of exe/cmd is required. The map key is the selector's role name
in status, metrics and alerts:
processes:
unifi: { cmd: "java .*unifi", user: unifi, group: unifi }
mongo: { exe: "${mongod_binary}", user: unifi }-
exe— exact resolved/proc/<pid>/exe(fail-safe; never cmdline). -
cmd— a Go RE2 regex matched against the process cmdline (argv joined). Use it for shared binaries (java .*unifi,openvpn .*tun1\.conf) when one executable serves several instances. The cmdline is spoofable, socmdnever authorizes signaling by itself — a kill still demands the exact resolved exe and the real UID — but it does narrow the identityforce_kill: autoderives, so a daemon that shares its executable with its own workload keeps that distinction when residuals are signalled. -
delegated—truemarks processes the service owns but Sermo must never signal: a workload tree the init unit deliberately keeps alive across a daemon restart (GlusterFS bricks and self-heal, container shims). They stay visible in monitoring, are never counted as residuals of a stop, and contribute no kill authority. Delegation flows down the process tree, because a workload process owns whatever it spawns: marking an SSH session covers the user's shell and everything that shell runs, none of which would match a selector of its own. Use it when the unit stops only its main process — systemdKillMode=process— so that stopping the daemon does not take its workload down with it.The SSH profile delegates authenticated sessions and connections still in the
[accepted]authentication phase. Detached user workloads such as atmuxserver can outlive their SSH parent; declare their verified executable and user withdelegated: truein the host override to preserve them across restarts. An undeclared orphan still blocks the following start. -
user/group— the process real UID / GID owner.
Do not use a generic helper executable shared by several units as a service selector. On systemd, cgroup attribution identifies the unit's processes; where there is no unique exact identity, leave the helper unselected. A broad helper selector can cross-attribute another unit's live process as a residual and safely block the restart.
When the init backend or a pidfile identifies a live process, discovery stays
within those attributed process trees (including all backend cgroup members).
Selectors label processes inside that scope instead of adding other instances
that happen to share an executable and user. This keeps per-instance memory,
CPU and replaced-binary notices separate. With no live backend or pidfile root,
selectors still discover processes across the host; shared executables need a
restrictive instance-specific cmd for that fallback.
These feed monitoring and the residual reaper, so a richer selector lets a
stop catch and kill more leftovers (an unkillable residual stays
orphan_processes). The process check still matches by exe/user only.
Set stop_policy.force_kill: auto to make every named selector that has both
an exact exe and user authorize cleanup of that same residual identity.
Sermo keeps each executable/user pair together — plus that selector's cmd, when
it declares one — sends TERM, rediscovers, then sends KILL only to the same
verified survivor. A selector with only cmd, an unresolved executable, a
delegated selector, or a process outside the configured identities remains
an orphan_processes failure and the restart never starts a second daemon.
force_kill: true still requires the explicit kill_only_if selector and is
the appropriate override when the configured process identities are not the
desired kill set; force_kill: false disables escalation.
Every start and stop Sermo issues is isolated from the init system's dependency graph, so restarting one service can never restart others. A restart is composed as stop then start, and both halves carry the flag — on systemd it is the stop that would otherwise drag down units bound to this one.
For a socket-activated systemd unit, leaving its socket untouched can make
systemd start that same unit again immediately after the isolated stop. When
the stop succeeded, every live residual is attributed by the backend to that
unit, and its state is again active, Sermo treats that as the completed restart:
it runs postflight but does not issue a second start or touch the activating
socket. This is not an exception for arbitrary leftovers: a selector-only
residual, an inactive/unknown unit, or any non-systemd backend still returns
orphan_processes and never starts the service.
| Backend | Command |
|---|---|
| systemd | systemctl <verb> --job-mode=ignore-dependencies -- <unit> |
| OpenRC | rc-service --nodeps <service> <verb> |
Only state-changing verbs are isolated. Status queries, reload and
reset-failed/zap never propagate, so they are issued unchanged.
This cuts both ways, deliberately. Isolation also means a start does not
pull up what the service requires: if a dependency is down, the start proceeds
and the service may fail on its own. systemd's documentation warns that
ignore-dependencies can leave the system inconsistent — that is the trade
accepted here, in exchange for never taking down a service nobody asked to
touch.
Measured on real hosts, nfs-server is the catalog exception that needs normal
dependency propagation. It ConsistsOf nfs-mountd and nfs-idmapd; an
isolated restart stops those companions but cannot pull the required
mount daemon up again, leaving NFSv3 mount requests unavailable while the kernel
NFS port remains healthy. The packaged nfs profile therefore sets
allow_dependencies: true. The NFS profile itself is a no-resident-process
service because its server runs in the kernel; the separate rpc-mountd profile
owns process discovery and stale-binary reporting for the userspace daemon.
Set the flag only on a service that is useless without the units it requires, and where you would rather it pull them up than fail:
name: some-service
allow_dependencies: trueOnly the packaged nfs service ships with it because the init graph is part of
that coordinating service's lifecycle. For ordinary companion units,
also_service: is the explicit form and remains preferable to relying on the
init system's graph.
It inherits from global defaults: like dry_run, so a whole host can opt back
in at once:
defaults:
allow_dependencies: trueDocker and libvirt services are unaffected: their backends have no dependency graph of this kind.
After a clean stop, the engine can verify the service left nothing behind:
stop_policy:
graceful_timeout: 30s
pidfile_absent: true # the declared pidfile must be gone
files_absent: # this instance's sockets/locks
- /run/postgresql/.s.PGSQL.5432
- /run/postgresql/.s.PGSQL.5432.lock
clean_after_stop: false # master opt-in: delete on stop- A lingering pidfile or
files_absentmatch is a warning (the stop still succeeds,ResultOK) folded into the result message and surfaced in CLI/web — it means the service crashed or left junk. Residual processes keep their strongerorphan_processes(red) handling via the reaper. clean_after_stopis the single master switch for all active deletion after a clean stop. It is opt-in (defaultfalse): with it off the engine only verifies and warns — it never deletes. Set it totrueto enable cleanup, which then does two things:- deletes any lingering
pidfile_absent/files_absentartifact (the oldrm-on-stop behavior), re-warning only if the delete fails; and - deletes the
clean_on_stoplist below.
- deletes any lingering
On Gentoo, PostgreSQL's init script can refuse to start after a crash because a
stale socket still exists. Once the instance's socket directory and port are
confirmed, opt into clean_after_stop with those exact paths so a Sermo restart
removes the stale socket after proving the processes have exited. Avoid a
wildcard spanning other PostgreSQL instances in the same directory.
clean_on_stop lists files and directories to delete on a clean stop (a
maintenance cleanup, distinct from the files_absent invariant). It only deletes
when clean_after_stop: true; listed without the master flag it is inert (so you
can stage the list and enable it later):
stop_policy:
clean_after_stop: true # required to actually delete
clean_on_stop:
- /run/svc/foo.tmp # a file
- /tmp/svc-*.lock # a glob (files)
- { path: /var/cache/svc, recursive: true } # a directory tree- A plain entry (string or glob) is deleted with
Remove(file or empty dir);{ path, recursive: true }deletes a directory tree (RemoveAll). - Safety (strict): every path must be absolute; a
recursiveentry must be a concrete (non-glob) path at least two levels deep and not the filesystem root or a shallow system directory (/,/etc,/usr,/var,/var/lib, …) — those are refused at validation time. A delete failure is a warning, not a failure.
A stopped service can leave an auxiliary process whose binary was replaced by
an upgrade. Both sermoctl start SERVICE and sermoctl restart SERVICE clean
up authorized residuals before starting, using the same stop_policy as stop.
force_kill: false still forbids that cleanup. A deleted executable needs the
additional kernel-file and ownership verification described in
safety.md; no extra YAML switch is needed.
Use sermoctl processes SERVICE to inspect the same deleted-executable candidates
that an operation sees. They retain their selector role and expose
exe_previous, external and signal_block_reason where applicable. An external
candidate that cannot be cleaned up blocks restart before stopping a healthy
current daemon. A helper is never evidence of a successful new main process.
The init backend attributes a whole control group to the service, so Sermo sees
processes there that no processes: selector claims. When such a process is also
outside the unit's principal process tree — reparented to PID 1 while still
counted against the unit — it is a stray: a probe that daemonized, a child the
daemon never reaped, a survivor of an earlier incarnation.
Strays are always reported (sermoctl processes shows stray=true). The optional
reap: block is what lets an operator clear them:
reap:
kill_only_if:
users: [root]
exe_any: [/usr/bin/dbus-daemon]- Without the block,
sermoctl reap SERVICE --applylists every stray and signals none. Authorization is opt-in per service and is never inherited fromdefaults:: one global selector would hand every service the same kill authority over processes none of them can name. kill_only_ifis the same paired selectorstop_policyuses and passes the same gate — exact resolvedexeand real UID; a delegated process never, an unresolvable exe never, PID 1 and kernel threads never.- Unknown keys under
reap:are rejected at validation time. A stray is reached by control-group membership rather than by a selector that named it, so a typo must not leave the action authorized by something you did not write. - No rule action can reap, and a stop never consults this block — a restart reports a stray it cannot clear rather than killing it. See safety.md for the full contract and cli.md for the command.
Detection is separate from clearing: Sermo injects a strays check into every
init-managed service that declares selectors, reporting the count and the
executables without alerting. See
configuration.md.
A catalog service can declare a top-level pidfile: <path> to wire both uses of a
pidfile from one line:
pidfile: /run/named/named.pidWhen a catalog service legitimately uses different pidfile names across distributions, declare candidates in preference order:
pidfile:
- /run/mysqld/mariadb.pid
- /run/mysqld/mysqld.pidWhen the pidfile is useful on one backend but legitimately absent on another (for example OpenRC writes one while a systemd unit runs the daemon in the foreground), keep the pidfile source for discovery but make the generated health check auxiliary:
pidfile: { path: /run/rngd.pid, optional: true }Use /run here, not /var/run. If a distro init script or service manager
reports /var/run/..., write the equivalent /run/... path in the catalog
service definition while preserving Linux/init compatibility. Before committing a new
pidfile or socket path, resolve it with readlink -f or inspect it with
namei -l; if any component is a symlink, use the resolved canonical target.
Sermo reads only the first line of a regular file there (at most 4 KiB); a FIFO,
device or directory at the pidfile path is reported as unreadable rather than
waited on.
On resolution this creates (a) an internal pidfile discovery selector — so the
parent process and its descendants are discovered and monitored without
adding a public processes: entry — and (b) a pidfile health check gated by
requires: [service]. Because of the gate, a missing or stale pidfile is
reported as an error only while the service is active (it means the service
died or lost its pidfile without the service manager noticing); a legitimately
stopped service is skipped, not alarmed.
A zombie PID counts as stale. When the service declares a processes: selector
with an exact exe, the check also fails if the live PID is neither named by
one of those selectors (a binary replaced on disk still counts) nor reported by
the init backend for the unit: after the daemon died, its PID may have been
recycled by an unrelated process. Without an exe selector (and in host
watches) the check can only verify that the PID is alive.
A check already named pidfile is respected, so a catalog service that needs a
custom check can still spell it out. Public processes: entries stay limited to
exe/cmd selectors with optional user/group; do not put pidfile under
processes:. The shorthand path can reference variables (e.g. pidfile: "${pidfile}") and accepts a scalar path, a candidate list, or {path: ..., optional: true}. Candidate lists are tried in order and pass on the first live
pidfile; if none exists, the backend PID fallback can still satisfy the gated
health check. optional: true keeps a missing pidfile as a warning instead of
making the service unhealthy.
When a single service owns several independent resident processes, use
pidfiles: as a map keyed by process role. Each role must also exist under
processes: with exact exe and user, so the pidfile PID can be tied back to
the process identity Sermo is allowed to observe:
pidfiles:
smbd: /run/samba/smbd.pid
nmbd: /run/samba/nmbd.pid
processes:
smbd:
exe: "${smbd_binary}"
user: root
nmbd:
exe: "${nmbd_binary}"
user: rootEach pidfiles.<role> creates its own internal pidfile selector and its own
gated health check (pidfile-smbd, pidfile-nmbd, ...). A value may still be a
candidate list for that specific role. Do not combine pidfile: and
pidfiles: in the same service: pidfile: means "one logical PID with
candidate paths"; pidfiles: means "all of these roles must have a live
pidfile."
A catalog service can declare a top-level Unix socket path when the active service should leave a socket behind:
variables:
socket: /run/cups/cups.sock
socket: { path: "${socket}", optional: true }On resolution this creates a socket health check gated by requires: [service]
and removes the top-level key. Like pidfile:, socket: accepts a scalar path,
a candidate list, or {path: ..., optional: true}. Use it for runtime sockets
owned by the service; protocol checks such as redis, dbus or libvirt still
use their own socket field inside the check body.
A catalog service can declare one regular lockfile created by the active service:
lockfile: /run/lock/subsys/smbOn resolution this creates a lockfile health check gated by
requires: [service] and removes the top-level key. Like socket:, lockfile:
accepts a scalar path, a candidate list, or {path: ..., optional: true}. It is
only evidence that the service left its own runtime lock artifact; it does not
block start/stop/restart/reload/resume and must not point under
<paths.runtime>/locks, which is reserved for Sermo operation locks.
Use an embedded type: dbus watch when a daemon owns a stable system-bus name.
The catalog profiles for systemd managers, NetworkManager, firewalld, TuneD,
GDM and several desktop/hardware daemons use the same check path as host
watches. Prefer the default peer probe when the object implements it; use
introspect to require a public interface, or property to read one stable
scalar property. Set require_owner: true for these resident daemons so an
active unit with a lost D-Bus registration fails instead of passing as merely
activatable. These probes disable D-Bus auto-activation and do not permit
arbitrary method calls. Adding a check-only watch does not add remediation;
attach a then: action only when that action has been reviewed independently.
See the D-Bus check reference.
The packaged catalog includes the resident Proxmox VE control-plane daemons, LXC/LXCFS, the Proxmox firewall variants, HA managers, API/SPICE proxies, QEMU event handling and ZFS ZED. It also includes commonly adjacent host daemons discovered on Proxmox nodes: Amazon SSM Agent, Cron, KSM tuning, the pNFS block mapper, and Prometheus IPMI/NUT exporters.
pve-container-%n is an instanced systemd profile. An active unit such as
pve-container@101.service materializes pve-container-101; inactive container
units are not inferred. Proxmox Perl daemons match the exact resolved Perl
executable and configured user before their narrow process title is considered.
The command line never authorizes signaling by itself. Exporter profiles verify
their local /metrics endpoint, while pvedaemon, pveproxy and spiceproxy
use TCP reachability because their HTTP endpoints intentionally require a
protocol-specific request or authentication.
These profiles add monitoring and normal init-backed operator controls only.
They do not ship automatic restart rules or enable SIGKILL. Boot setup units,
systemd infrastructure, filesystem mount helpers and hardware state belong to
their existing host watches rather than duplicate catalog services. Debian's
dm-event and smartmontools unit names are aliases of the existing dmeventd
and smartd profiles.
Some applications ship one binary per version and several can be installed at
once (php-fpm, postgres, tomcat, erlang/beam, berkeley db). Instead of one file
per version, write a single app version template whose name: contains
%v, with ${version} in the discovery path. A service template with the same
token links that app.
name: postgres-%v
display_name: "PostgreSQL ${version}"
variables:
binary: "/usr/lib64/postgresql-${version}/bin/postgres"
preflight:
binary: { type: binary, path: "${binary}" }
version: { type: command, command: ["${binary}", "--version"], timeout: 10s }
---
name: postgres-%v
display_name: "PostgreSQL ${version}"
service:
systemd: ["postgresql-${version}", "postgres-${version}"]
openrc: ["postgresql-${version}", "postgres-${version}"]
apps: ["postgres-${version}"]
variables:
data_dir: /var/lib/postgresql/${version}/data
pidfile: "${data_dir}/postmaster.pid"On load, Sermo discovers app versions by globbing the linked app's
variables.binary path with ${version} wildcarded (here
/usr/lib64/postgresql-*/bin/postgres) and extracting what filled it. Service
templates in catalog/services prefer the active init service as source of
truth: token-bearing service: candidates are matched against active
systemd/OpenRC units, and only matching services materialize. Each match becomes a
concrete app or service with %v and ${version} substituted everywhere (name,
display_name, service, app links, ...) — postgres-14, postgres-16, ... — and
the templates themselves are dropped. If nothing is installed or no matching
service is active, the template yields nothing. The YAML filename does not have
to match name:; keep one descriptive file for the template and treat name:
as the catalog identifier. %v may sit anywhere in the name (db%vsql →
db4.8sql). Note: %v is substituted only in the name; inside the body always
use ${version} (e.g. in service or apps).
Prefer application discovery in catalog/apps when the installed binary path
identifies the version or instance. A versioned or instanced service that links a
matching app, such as apps: ["postgres-${version}"] or
apps: ["php-fpm${version}"], uses that app for runtime binary validation. For
catalog services, put the same tokens in service: so the service materializes
from the unit that is actually active on the selected init backend.
variables.binary may be a string or a candidate list. Use it when the
versioned path is also the runtime executable that preflight and version checks
should probe. For app and library templates that discover from versions.from
and do not declare variables.binary, the materialized document binds
${binary} to the path that matched; keep versions.from for discovery sources
that are not the runtime executable.
When an app or library cannot discover from its runtime executable, use
versions.from there and link the generic or versioned app that owns the binary:
name: myservice-%i
versions:
from: "/etc/myservice/${instance}.conf"
variables:
binary: /usr/sbin/myservice
preflight:
binary: { type: binary, path: "${binary}" }versions.from is discovery-only metadata; it never appears in materialized apps
or services. Its path must contain every value marker in the template name
(${version}, ${n} and/or ${instance}; ${sep} is optional); otherwise it is
ignored and apps/libraries fall back to variables.binary. Matches are de-duplicated
by their materialized token tuple.
A discovered version must start with a digit, so siblings of an unbounded
trailing placeholder (a bare php-fpm symlink, a php-fpm.conf) are not mistaken
for versions. Even so, a placeholder bounded on both sides (e.g.
/usr/lib64/php${version}/bin/php-fpm, in the app variables.binary path) discovers most
precisely.
Some packages ship one binary per subcommand and no plain versioned entry point:
Berkeley DB installs db5.3_archive, db5.3_dump, db5.3_stat, … and no
db5.3. A trailing ${version} captures everything after the name, so each
subcommand would materialize its own app (Berkeley DB 5.3_archive,
Berkeley DB 5.3_dump, …).
versions.suffix names the part of the captured value that is not the version.
It takes one glob or a list of them, anchored at the end of the value; the
longest match is trimmed, and every trimmed value then de-duplicates into a
single instance:
name: db%v
display_name: "Berkeley DB ${version}"
versions:
from: ${bindir}/db${version}
suffix: "_*"
variables:
binary:
- ${bindir}/db${version}_dump
- ${bindir}/db${version}_stat
preflight:
binary: { type: binary, path: "${binary}" }
version: { type: command, command: ["${binary}", "-V"], timeout: 10s }db5.3_archive, db5.3_dump and db5.3_stat all trim to 5.3, so one
db5.3 app is registered per installed release. A value the suffix does not
match, or would consume entirely, is kept whole — a bare db6.2 still
registers as 6.2. A suffix must begin with a literal separator; a leading *
or ? is rejected because it would swallow the version itself.
Pin variables.binary when the family is discovered this way. Discovery from
versions.from leaves the declared binary alone, so the candidate list decides
which subcommand preflight probes instead of whichever one globbed first — it
matters when some of them behave differently (db5.3_tuner rejects -V).
%v/${version} accepts a digit-leading version (8.3, 12.0.2); use
%n/${n} when the value is a plain integer — it matches only whole
numbers, otherwise working exactly like %v:
name: python%n
display_name: "Python ${n}"
variables:
binary: "/usr/bin/python${n}"
preflight:
binary: { type: binary, path: "${binary}" }/usr/bin/python* then materializes python2/python3, but not python3.11 or
python-config.
When a simple %v or %n template also has an unversioned active-slot binary,
Sermo materializes it automatically. If /usr/bin/python exists, this registers
python in addition to python2/python3; when it is absent, only the numbered
binaries are registered. The empty token is substituted before name,
display_name and description are trimmed, so display_name: "Python ${n}"
becomes Python for the active slot. Composite templates (%i plus %v, a
separator token, etc.) do not infer that entry from versions.from; declare
versions.current_from when they have a concrete active-slot executable such as
/usr/bin/java. That path materializes the unversioned base name before the
first token (java-%i-%v -> java) and becomes its ${binary} when the
template does not declare one. current_from may also be a list of direct paths:
versions:
current_from: /usr/bin/javaSet versions.unversioned: false to ignore the marker-less or current_from
active slot; a map form can still override fields for the unversioned instance
when a template needs a custom label:
name: python%n
display_name: "Python ${n}"
versions:
unversioned:
description: "Active Python interpreter"
variables:
binary: "/usr/bin/python${n}"
preflight:
binary: { type: binary, path: "${binary}" }An app or library template materializes one instance per real binary and
version. When overlapping binary: candidates match the same file and version
with different token values — /usr/lib/jvm/openjdk-bin-17 read as instance openjdk by
${instance}-bin-${version} and as openjdk-bin by ${instance}-${version} —
the first candidate in the list wins, so list the most specific pattern first.
Different versions sharing one wrapper (Gentoo links python2 and python3 to
the same python-exec) stay separate. The unversioned active slot is kept alongside its versioned instance, and catalog
service templates are not affected: their instances are init units, and several
of them (PHP-FPM pools, Tomcat instances) share one binary.
If a template would materialize a name: that already exists as an explicit
document in the same catalog category, validation reports a collision. Remove
one definition or adjust the template discovery; Sermo does not silently choose
between an explicit document and a generated one.
Templates may also use ${current} in display_name or description. During
materialization it becomes current only for the versioned entry whose binary is
the same filesystem entry as the active-slot binary, whether discovered from the
marker-less path or declared with versions.current_from (for example
/usr/bin/php -> /usr/bin/php8.2 or /usr/bin/java pointing at the active JVM);
otherwise it becomes empty before metadata is trimmed. This lets
display_name: "PHP ${version} ${current}" render as PHP 8.2 current for the
active version and PHP 8.3 for the others without running version commands
during config load. Symlinks are resolved before comparison. App/service
inventory commands may still add the current label at inspection time when an
active-slot wrapper reports the same version_short as one materialized
version, which keeps wrappers such as Gentoo Java generic without from_file
catalog metadata.
Use %i/${instance} for named init instances discovered from bounded service
metadata. Scope backend-specific discovery to matching service candidates; for
example, an OpenRC-specific profile can expose only service.openrc: ["openvpn.${instance}"], while a systemd template can expose
service.systemd: ["openvpn-client@${instance}"].
Some services encode both a version and an environment/pool in one name, joined
by - or _ — tomcat-8.5-main, tomcat-9-guacamole, php-fpm8.4_airbnb. Use
%s/${sep} for that joining separator, which matches an empty string, - or
_. A name may carry several tokens (tomcat-%v%s%i); for service templates they
are discovered together from active service units whose service: candidates
contain the same markers, and bound everywhere at once. A non-final %v is
bounded so it stops at the separator (8.5), and the instance may be empty —
when it is, the separator collapses too, so a bare tomcat@8.5.service
materializes tomcat-8.5 with no trailing -:
name: tomcat-%v%s%i
service:
openrc: ["tomcat-${version}${sep}${instance}"]
systemd: ["tomcat@${version}${sep}${instance}"]A service template in catalog/services normally discovers from active init
units. Put every supported service spelling in service: and split it by backend
when systemd/OpenRC names differ. The linked app (generic like openvpn, or
versioned like php-fpm${version}) still supplies ${binary} for preflight and
process identity. A service never discovers from its own binary.
When a generic unit spelling could also name a different daemon, set
versions.require to a path containing the same template variables. Sermo
materializes a single-token or composite instance only if at least one required
path exists; this keeps overlapping unit names out of the wrong catalog profile.
When discovery comes from init service metadata, let the linked app own runtime
binary validation when it is versioned. For example, PHP-FPM links
php-fpm${version}; that app already validates /usr/sbin/php-fpm${version} or
/usr/bin/php-fpm${version}, so the service does not repeat the same candidates
in versions.require:
service:
systemd:
- "php-fpm@${version}${sep}${instance}"
- "php-fpm@php${version}${sep}${instance}"
- "php-fpm-php${version}${sep}${instance}"
- "php${version}${sep}${instance}-fpm"
- "php-fpm${version}"
openrc:
- "php-fpm-php${version}${sep}${instance}"
- "php${version}${sep}${instance}"
- "php-fpm${version}${sep}${instance}"
- "php-fpm${version}"
apps: ["php-fpm${version}"]
pidfile:
- "/run/php-fpm/php-fpm-${version}${sep}${instance}.pid"
- "/run/php-fpm/php-fpm-php${version}${sep}${instance}.pid"
- "/run/php-fpm-php${version}${sep}${instance}.pid"
watches:
pidfile:
check:
type: pidfile
optional: true
path:
- "/run/php-fpm/php-fpm-${version}${sep}${instance}.pid"
- "/run/php-fpm/php-fpm-php${version}${sep}${instance}.pid"
- "/run/php-fpm-php${version}${sep}${instance}.pid"
requires: [service]Put the exact systemd instance first in service.systemd, e.g.
php-fpm@${version}${sep}${instance} for php-fpm@8.2.service. Avoid a generic
php-fpm systemd fallback in versioned templates: it can make several
discovered PHP-FPM versions operate on the same unit. The pidfile check is
optional because some systemd units publish MainPID even when the declared
PIDFile= is not written.
PHP-FPM's worker selector and residual cleanup policy use variables.user:
apache on Gentoo and www-data otherwise. Set this variable to the pool's
actual user when it differs. After a master crash, workers can survive without
a usable pidfile; their exact executable and user must still match before
Sermo can clean them up and start the service.
An entry under processes, watches or preflight may carry an
enable_if guard that keeps it only when a key in a distro config file satisfies
a predicate; otherwise the entry is dropped during service resolution. This
models components that are optional per host — e.g. a Samba profile monitors
winbindd only when /etc/conf.d/samba's daemon_list names it. Do not link
such a component under apps, because linked apps are mandatory preflight
dependencies for service operations:
processes:
winbindd:
exe: ${winbindd_binary}
enable_if:
file: /etc/conf.d/samba
key: daemon_list
contains: winbindd # or: equals: <value> | matches: <regex>
watches:
winbindd:
enable_if:
file: /etc/conf.d/samba
key: daemon_list
contains: winbindd
check:
type: process
exe: ${winbindd_binary}
state: runningThe key is read from KEY="val", key: val, key = val and whitespace-only
key val lines alike (OpenRC and dnsmasq, YAML, exim.conf, snmpd.conf); the
first line that starts with the key wins, a key alone on its line is a flag
with an empty value, and a trailing comment stays part of the value.
A missing file or absent key prunes the entry (fail-safe). The guard is stripped
from surviving entries. A host watch document (under paths.watches) accepts
the same top-level enable_if; the daemon then skips that watch on hosts where
the gate fails. config validate still checks disabled entries before
they are pruned, so typos in optional process/check definitions are reported.
enable_if is intentionally not supported under rules, policy, guards or
other safety-affecting sections.
Instead of a config-file predicate, enable_if may name the init backend the
entry belongs to: enable_if: { init: openrc } (or systemd) keeps the entry
only when that backend is active, and excludes the file/key predicate form.
Use it for components that exist under one init system only. The packaged
salt-minion profile gates its Gentoo supervise-daemon selector this way:
under OpenRC the supervisor is the one strict identity Sermo may signal, while
on Gentoo running systemd no supervisor exists, and an exact selector that can
never match a live process would make the restart identity guard block every
operation on the unit:
processes:
supervisor:
exe: /usr/bin/supervise-daemon
user: root
enable_if: { init: openrc }The config reader accepts key=value assignments, YAML key: value block
mappings and bare key feature flags.
Use equals: "" to match a bare flag. The packaged dnsmasq profile uses this
to add its DHCP check only when /etc/dnsmasq.conf has dhcp-range=..., and its
TFTP check only when that file has enable-tftp. If those directives live in a
file under /etc/dnsmasq.d, override the corresponding
watches.dhcp.enable_if.file or watches.tftp.enable_if.file in the configured
service. Override dhcp_host/dhcp_port or
tftp_host/tftp_port/tftp_query when the local endpoint differs.
The packaged cloudflared profile is the YAML case: cloudflared tunnel ingress validate only validates locally declared ingress rules, so it fails on a
remotely-managed tunnel whose config.yml carries just a token: — and a failed
preflight blocks every operation, restart included. Its config preflight is
therefore gated on ingress being present in the file.
A variable may take its value from a config file instead of a literal, useful when
a port or path is defined in the service's own config. directive: reads the token
after a key value line (OpenVPN/sshd style); pattern: reads capture group 1 of
a regex; default: applies when the file or key is absent:
variables:
config: "/etc/openvpn/${instance}.conf"
port:
from_file: "${config}"
directive: port # "port 1194" -> 1194
default: 1194 # required fallback when file/key is absent
# tomcat: pattern: '<Connector[^>]*?\bport="(\d+)"'It is evaluated during resolution (so it can reference other variables such as
${config}) and re-evaluated on every config reload. A from_file path or
pattern may also name another from_file variable, for example an include
file read from the main config; such variables are read in dependency order,
so each sees the file value of the one it names, and a reference cycle is a
validation error. pattern may also
reference variables such as ${instance}; those values are escaped as regex
literals before the file is read. The variable spec must define from_file,
default, and exactly one of directive or pattern. pattern must compile
and include a capture group. A missing file or unmatched key uses default;
malformed specs or unknown variables in from_file / pattern are validation
errors.
sermoctl apps reports the applications described by catalog apps: which are
installed (their binary is present and executable), whether their health
command succeeds when configured, and the version their version command
reports. The VERSION column shows the short version by default; add --long to
show the full raw string.
APPLICATION VERSION STATUS
Nginx 1.24.0 ok
Python 3 3.11.2 ok
Redis - error: /usr/bin/redis-server is not executable
$ sermoctl apps --long
APPLICATION VERSION STATUS
Nginx nginx version: nginx/1.24.0 ok
Python 3 Python 3.11.2 ok
Only installed applications are shown; sermoctl apps all also lists the rest as
not installed. The same --long and all apply to sermoctl libs and
sermoctl services catalog. With version templates this lists each installed version as
its own row (e.g. PHP-FPM 8.3, PHP-FPM 7.4). For sermoctl services catalog, version
commands are best-effort inventory data: a failed distro-specific version probe
leaves the version unknown instead of marking the installed service as an error.
--json is unaffected by --long — it always emits both, with the structured
name, display_name, binary, version, version_short,
version_source, installed, ok and status.
When an app declares health, Sermo uses it as the preferred health probe for
sermoctl apps/libs/services catalog and the WebUI application list. Only the exit
code is evaluated (expect_exit, default 0, or a list such as [0, 1]);
stdout/stderr matchers and the printed output are ignored for health. The
version command is only used as a fallback health probe when no health
command exists; when health exists, version reports display data and a
version failure does not override health.
Do not mark an app version probe optional unless the app also has a health
probe; otherwise Sermo can only prove that the binary exists, not that it can run.
For catalog apps that are separate binaries from the same package, version_from
can point at another catalog app whose version probe supplies the displayed
version. The app still checks its own variables.binary and health;
version_from only
sets version/version_short when the app has no local version result.
Catalog apps can use version_match when a binary name is shared by compatible
implementations. It runs against the combined stdout/stderr of the local
version command and supports contains, excludes and regex. If it fails,
the app is treated as not installed rather than as an installed app with a bad
version. For example, MariaDB accepts mysqld only when the output contains
MariaDB, while MySQL excludes that token so MariaDB's compatibility mysqld
does not appear as MySQL.
version is the raw first line the version command prints (e.g. nginx version: nginx/1.30.2); version_short reduces it to just the numeric version and at
most the patchlevel (1.30.2), taking the first major.minor[.patch] token and
dropping any further build components and suffixes (so 2.8.4.1-0+g… becomes
2.8.4 and 4.2.8p18 becomes 4.2.8). If there is no dotted token, a guarded
integer-only version N token is accepted for projects such as polkit and
date-coded numad releases. It is empty when the version line carries no
recognizable number.
A catalog service may instead declare a dedicated version_short command (under
preflight or commands, alongside version) that prints the bare version
itself, sidestepping the regex when a tool can report it directly. Its first
non-empty output line is then used verbatim. The packaged interpreter apps do
this with their resolved binary — e.g. PHP runs php -r 'echo PHP_VERSION;',
Python runs python -c 'import platform;print(platform.python_version())', Node
node -p process.versions.node — so their short version never depends on
parsing. When no such command is configured (or it errors or prints nothing),
version_short falls back to parsing the version line as above.
preflight:
health: { type: command, command: ["${binary}","-h"], timeout: 10s }
version: { type: command, command: ["${binary}","-v"], timeout: 10s }
version_short: { type: command, command: ["${binary}","-r","echo PHP_VERSION;"], timeout: 10s }A service template may uses a base catalog service to inherit its checks,
processes and rules, while a linked app supplies the instance- or
version-specific binary:
name: myvpn-%i
uses: myvpn # an existing catalog service, not a template
display_name: "MyVPN ${instance}"
service: { systemd: ["myvpn@${instance}"], openrc: ["myvpn.${instance}"] }
apps: ["myvpn-${instance}"]Only a template's uses is followed, and only one level deep. Validation
rejects a template whose uses names a missing catalog service or another
template, and uses on a catalog service that is not a template, since
nothing would inherit it.
A configured service then targets a concrete instance: the packaged nebula-%i
template materializes nebula-nebula0 from nebula@nebula0.service, and a
service file selects it with uses: nebula-nebula0.
Active systemd/OpenRC units normally materialize catalog instances for discovery.
An explicitly configured uses: instance also materializes when its unit is
stopped or failed, so sermod can report that service state instead of rejecting
the whole configuration.
Nebula Mesh's accompanying components are cataloged separately as
nebula-agent and nebula-mgmt. Both resolve their native systemd or OpenRC
unit and exact process identity. The management-server profile reads listen:
from /etc/nebula-mgmt/server.yml (falling back to 127.0.0.1:8080) and checks
its unauthenticated /readyz endpoint. These profiles are monitor-only: they
do not add automatic restart rules for certificate or control-plane services.
The service's identity is its name; service declares the init-unit
name(s) to operate on. The simplest form is a single name that works on both
init systems:
service: apache2When the unit name differs across init systems, list per-init candidates; Sermo
resolves the first one the active backend actually knows (systemd via
systemctl cat, OpenRC via the init script):
service:
systemd: [apache2, httpd]
openrc: [apache2, apache]Candidates are bare names — systemd appends .service automatically. They are
tried in order and deduplicated, and the resolved name is used for all later
operations. A scalar service is trusted even when the probe cannot surface
it (e.g. sysv-generated units). A per-init list first requires a backend
match; if the probe cannot surface one, Sermo logs or prints a warning and falls
back to the configured seed unit so sermod, the web UI and sermoctl behave
the same on historic init-service setups. An init system with no entry means the
service is not available there: sermod skips it and reports the skip as an
informational notice (not a warning), because the map itself declares the
backend unsupported. Services using control: (libvirt/docker) do not use the
init-unit fallback.
An enabled instance can override the unit with a scalar (e.g.
service: redis-cache) to run as its own unit, or omit service entirely to
inherit the catalog service's candidates.
A service may clone another service to make a second instance:
name: redis-cache
clone: redis-main
variables:
port: 6380
pidfile: /run/redis-cache/redis.pidClone copies the source before variable expansion, so overriding the port
variable alone is enough — every check that references ${port} resolves to the
new value. Clone chains resolve transitively; cycles are rejected.
To run several instances of the same application — same binary, same checks and
rules, different listen port, pidfile and config file — let each instance uses
the catalog service and override only its unique variables.
The catalog service parametrizes everything that varies with ${...} placeholders and
threads each one into the commands and checks that consume it. In particular the
config-file path should be a variable wired into every command that reads it, so
two instances never pick up each other's configuration:
name: dbserver
variables:
port: 3306
pidfile: /run/dbserver/main.pid
config: /etc/dbserver/main.cnf
pidfile: "${pidfile}"
watches:
tcp:
check: { type: tcp, port: "${port}" }
config:
check: { type: command, command: ["dbserverd", "--defaults-file=${config}", "--help"] }Each instance overrides the three variables and gives itself an init unit (a
systemd template instance or a distinct unit name) with a scalar service:
name: db-inst1
uses: dbserver
service: db-inst1
variables:
port: 3306
pidfile: /run/dbserver/inst1.pid
config: /etc/dbserver/inst1.cnfA second instance is the same file with its own name/unit and variables (e.g.
name: db-inst2, service: db-inst2, port: 3307, the inst2.* paths).
Prefer uses over clone here: every instance derives from the
catalog service and only overrides variables. Reach for clone only when one instance
should copy another concrete service almost verbatim. See docs/sermo-all.yml
for a complete worked configuration.
Network probes in a profile use its variables.host and variables.port;
override these on the service instance when the listener differs from the
catalog default. Protocol-specific ports, such as Dovecot's pop_port, remain
separate variables.
Catalog alerts are graded (see Severity) so each
notifier's min_severity can filter them. The resource alerts are advisories
that escalate: alert-if-memory-high warns above 30 % of host RAM and becomes
an error above 50 %, alert-if-cpu-thread-high warns above 90 % of one core and
becomes an error above 98 %. A stopped service or missing daemon is critical;
a certificate inside its renewal window is a warning and an expired one is
critical. Override the base threshold and the level together on an instance
that needs different numbers, for example
levels: { error: { op: ">", value: 70% } } beside value: 50%.
MySQL and MariaDB validate the selected variables.config with
--defaults-file=... --help --verbose. The defaults-file argument comes first
and a missing or invalid file fails the preflight. This is a compatible option
parser check, not a full startup or data-integrity check. High service memory
usage only alerts; memory_alert_threshold defaults to 80% of host RAM and
should reflect the instance's buffer-pool budget. The alert is a warning and
escalates to error past memory_error_threshold (90%); raise both
together. The separate memory check
still reports usage over 60%. MySQL, MariaDB, PostgreSQL and Backrest backup
guards block both stop and restart; Sermo named locks remain the preferred
way to protect jobs that can be wrapped with sermoctl lock.
Grafana's health check requires HTTP 200 and JSON database: ok. Prometheus
also checks /-/ready with HTTP 200, so a live process still loading its TSDB
does not count as ready. Both checks verify startup and alert after three
failed cycles; neither adds automatic restarts.
Redis and KeyDB provide an opt-in AOF write check. Enable it only on instances
with appendonly yes: servers without AOF may omit aof_last_write_status.
The monitoring user must be allowed to run PING and INFO. Missing fields,
denied INFO access and a write error fail the enabled check.
name: redis-main
uses: redis # keydb supports the same watch
watches:
alert-if-aof-write-failed:
enabled: trueRedis and KeyDB also provide an opt-in alert-if-maxmemory-near-limit watch:
a warning when used_memory stays at or above variables.maxmemory_used_limit
percent of maxmemory (default 90) for five minutes. Enable it where reaching
the limit hurts — under maxmemory-policy noeviction a full server rejects
writes while the process, the port and PING all stay healthy, and a session
store under an LRU policy silently evicts live sessions. Leave it disabled on a
pure LRU cache, which sits at its limit by design.
name: redis-main
uses: redis # keydb supports the same watch
variables:
maxmemory_used_limit: 85
watches:
alert-if-maxmemory-near-limit:
enabled: trueBIND's named profile keeps one liveness probe (port: an A query for
variables.query, default localhost) and offers two opt-in watches for a
recursive resolver. recursion asks variables.host to resolve
variables.recursion_query (default example.com) and fails on REFUSED
(an allow-recursion/allow-query ACL that leaves out the probed address) or
SERVFAIL (upstream unreachable, DNSSEC validation failure), with the rcode in
the watch data; the liveness probe alone keeps passing because the daemon is
up. Declare one such watch per listener address a client population uses
(LAN, VPN) so an ACL that covers only some networks is caught. resolver
resolves the same name through the host's own /etc/resolv.conf
(resolvconf: true) and belongs on a host whose resolver is this server with
no fallback: when named stops answering, the host itself loses name
resolution, including the one Sermo needs to deliver notifications.
name: named
uses: named
watches:
recursion:
enabled: true
recursion-vpn: # a second listener address with its own ACL
check:
type: dns
host: 10.200.200.1
query: example.com
timeout: 3s
expect: { rcode: NOERROR, answers: { op: ">", value: 0 } }
for: { cycles: 2 }
resolver:
enabled: true
zone-authoritative: # the zone the server is primary for is loaded
check:
type: dns
host: 127.0.0.1
query: example.internal
qtype: SOA
expect: { rcode: NOERROR, aa: true }A log watch on named's own log catches what a probe from one address cannot:
security: info: client ... denied (allow- lines name every client an ACL
rejected, and lame-servers: info: (broken trust chain|no valid RRSIG|RRSIG failed to verify) lines are real DNSSEC validation failures (each one a
SERVFAIL to a client). The dnssec: info: validating ...: no valid signature found lines are not: a validating resolver logs them for every unsigned
delegation it proves insecure and still answers NOERROR.
PHP-FPM's fpm check compares the current listen_queue with
variables.listen_queue_max (default 0). Add a sustained alert after
configuring the pool's ping.path and pm.status_path; the rule reuses the
existing check. The cumulative max_children_reached counter is informational,
since an old peak alone does not indicate current saturation.
name: php-fpm8.4
uses: php-fpm8.4
variables:
status_path: /status
listen_queue_max: 5
watches:
fpm:
check:
socket: /run/php/php8.4-fpm.sock
optional: false
rules:
alert-if-listen-queue-high:
type: alert
if:
failed: { check: fpm }
for: { duration: 2m }
then:
action: alert
message: PHP-FPM ping/status is unavailable or its listen queue is highThe catalog leaves this rule opt-in so installations that disable the fpm
watch have no dangling check reference. Remove the rule when disabling that
watch. The rule alerts after two minutes of failed checks, including unavailable
ping/status endpoints. No pool configuration is changed by Sermo. For deeper
database checks, use an account restricted to monitoring; see
mysql-query-health.yml for an
authenticated SELECT 1 check. The default credential-free MySQL/MariaDB probe
only reads the server greeting.
watches:
http:
enabled: false # keep but disable
ping:
delete: true # remove the inherited entryThe top-level monitor flag sets a service's monitoring behavior when the
daemon starts:
name: web
uses: nginx
monitor: enabled # enabled (default) | disabled | previousenabled(the default when the flag is absent): always monitor on startup.disabled: never monitor — the worker exists but every cycle is skipped.previous: restore the runtime state the service had before the daemon last stopped. On the very first run (no recorded state) it defaults to monitored.
Top-level enabled: false disables the service entirely; no worker is built.
With monitor, the worker exists and only check/rule execution changes.
The live state is toggled at runtime with sermoctl monitor <svc> /
sermoctl unmonitor <svc> and persisted in the state database under
paths.state (see configuration). Because that database
survives reboots, a previous service comes back up in whatever state an
operator last left it.
Host watch documents use the same top-level
monitor: enabled | disabled | previous values; see
configuration.
A service may also carry its own watches: block — per-service watches that can
fire a hook/notification or compact then.action, and can use the service-scoped
service/metric/process_count check types. See
Service watches.
Host-global checks such as terminal_sessions also work there: place a
tmux or screen check under the SSH service to show its
configured users' terminal sessions in the Web UI Sessions panel. It does
not create a systemd/OpenRC service. Administrators may close one exact
displayed session; Sermo freshly revalidates its multiplexer-owned generation
and uses the configured client as that user, without signalling a guessed PID.
The windows inside a screen or tmux server are terminals of that session, not
SSH sessions: the SSH source leaves them out even though screen records them
in utmp with a remote-looking host, and closing the multiplexer session closes
them.
The remote deployment generator creates these checks for active multiplexer
namespaces that its read-only inventory can attribute to a local user; named
tmux sockets are kept separate, while inactive or unattributable namespaces
are omitted.
The packaged fcron profile treats more than one process in the service tree as
an active scheduled job. While that condition is active, its guard blocks
restart and stop, allowing the job to finish before an operator retries the
operation.
Restarting a database, cache or FTP server with active clients is a site policy,
not a catalog default. Add one of the opt-in examples to the concrete service
you enable: MySQL,
PostgreSQL,
Redis, or
ProFTPD. The
database examples count application sessions with read-only SQL; Redis uses its
native connected_clients metric; FTP uses tcp_connections, which counts
control-channel TCP sockets rather than authenticated users. See
connection guards for the safety behavior.
The mysql, mariadb and postgres catalog services ship the
alert-if-query-long-running watch, a db_queries
sensor: one warning per statement that runs longer than
long_query_duration (default 5m), with the statement text, user, database
and client host, recovered when the statement ends. Each statement is its own
event_notify incident. The dashboard's Sessions panel and
sermoctl sessions SERVICE list the running statements.
| service | variable | default | meaning |
|---|---|---|---|
mysql, mariadb |
db_socket |
/run/mysqld/mysqld.sock |
Unix socket the watch connects to |
mysql, mariadb |
defaults_file |
/root/.my.cnf |
option file whose [client] credentials it uses |
postgres |
host, port, monitor_user, database |
127.0.0.1, 5432, postgres, postgres |
connection, shared with the replication watches |
| all three | long_query_duration |
5m |
the alert threshold |
Credentials. On MySQL/MariaDB the watch connects the way mysql run by
root does: over db_socket, with the user/password of the [client]
(also [client-server], [client-mariadb], [mysql]) groups of
defaults_file. A missing file just means no password and the user defaults
to root, which works with the unix_socket authentication most
distributions give root. To list other accounts' statements the account needs
PROCESS; to cancel them from the Sessions panel it needs CONNECTION_ADMIN
(MySQL 8) or SUPER, or it can only cancel its own. On PostgreSQL
monitor_user needs pg_monitor to read other roles' statement text (the
default postgres superuser has it) and pg_signal_backend (or ownership of
the backend) to cancel a statement.
Tuning per host — a services.local override, without editing the service
file:
# /etc/sermo/services.local/mariadb.yml — reporting server: long queries are normal
name: mariadb
variables:
long_query_duration: 30m
watches:
alert-if-query-long-running:
check:
exclude_users: [backup, etl]# /etc/sermo/services.local/mysql.yml — no statement alerts on this host
name: mysql
watches:
alert-if-query-long-running:
enabled: falseThe catalog never ships an automatic kill. To cancel statements automatically,
add an explicit, scoped then.kill_query with its own policy: to the
concrete service — see killing a statement
and safety.
The postgres catalog service ships five replication sensors built on the sql
check: alert-if-replication-slot-backlog, alert-if-logical-slot-unconfirmed,
alert-if-replication-slot-inactive, alert-if-replication-replay-lag and
alert-if-standby-replay-delay. They are tuned by these variables:
| variable | default | meaning |
|---|---|---|
monitor_user |
postgres |
role the replication queries connect as |
database |
postgres |
database the queries connect to |
slot_backlog_mib |
1024 |
WAL retained by the most lagging slot, in MiB |
logical_unconfirmed_mib |
512 |
data a logical consumer has not confirmed, in MiB |
replay_lag_mib |
256 |
sent but not replayed by the most lagging replica, in MiB |
standby_delay_seconds |
300 |
how far behind a standby may fall, in seconds |
The thresholds are plain numbers because a sql check compares numerically and
does not take size suffixes, so the queries return MiB and seconds directly.
Pick them well above the idle baseline: a healthy primary already retains around
16 MiB (one WAL segment) because a logical slot's restart_lsn trails by design.
monitor_user must hold pg_monitor (the default postgres superuser
does). A role without pg_monitor or pg_read_all_stats still sees rows in
pg_stat_replication, but with sent_lsn/replay_lsn as NULL for other
backends — the aggregate then collapses to 0 and the lag watches never fire,
silently. pg_replication_slots is readable by any role, so the slot watches
are unaffected.
These watches only make sense where replication actually happens. Enable only the watches that match the PostgreSQL role and replication features on that host.
The exim catalog service exposes separate confirmed operator buttons for
exim_tidydb on the callout and retry hints databases. Its matching
service watches count the records, publish records metric series and can run
the same bounded cleanup command when the configured limit is exceeded.
Exim 4.99 and newer name the SQLite record table tblblob; older supported
releases use tbl. The catalog defaults callout_db_table and
retry_db_table to tblblob. Fleet installation discovers each database's
schema read-only and overrides either variable with tbl when required. A
non-SQLite database, absent file or unsupported schema disables only the
affected record watch; exim_tidydb buttons remain available because that
utility supports Exim's native hints backend independently of the graph query.
The exim catalog service alerts on a mass mailing (a stolen password, a
looping application, a spam run) from three angles, each tunable through a
variable:
| Watch | Signal | Warning variable (default) | Error variable (default) |
|---|---|---|---|
alert-if-queue-high |
exim -bpc above the limit for 3 minutes |
queue_limit (200) |
queue_error_limit (1000) |
alert-if-msglog-backlog-high |
files under msglog_dir, counted recursively |
msglog_backlog_limit (200) |
msglog_backlog_error_limit (1000) |
alert-if-msglog-backlog-growing-fast |
msglog growth inside msglog_growth_window |
msglog_growth_limit (100), msglog_growth_window (2m) |
— |
alert-if-memory-high |
resident memory of the Exim processes | memory_limit_bytes (104857600, 100 MiB) |
memory_error_limit_bytes (524288000, 500 MiB) |
Each alert is a warning at its first variable and
escalates to error past the second, so a small
backlog reaches the chat channel while a mass mailing also reaches a notifier
with min_severity: error.
The queue watch next to them is graph-only: it publishes the queue depth as
a messages series and never alerts. The msglog counts are recursive because
split_spool_directory (the usual production setting) spreads the files over
62 hashed subdirectories, where a flat count reads zero. The memory ceiling is
absolute rather than a host percentage because Exim idles at a few tens of MB
and a mailing that pushes it to 3 GB is still under 5% of a 64 GB host.
Raise queue_limit on a relay that legitimately holds a deep queue, and raise
queue_error_limit with it: an error level that is not above the warning
threshold is ignored (sermoctl config validate warns about it). A
service-level override keeps the rest of the profile:
name: exim
uses: exim
variables:
queue_limit: "2000"
queue_error_limit: "10000"The alloy catalog service watches the collector from the outside, the way
its clients see it, and restarts it when it stops serving them. It was
reshaped after an Alloy built with Go 1.26 leaked one MPTCP socket per
accepted OTLP connection on a kernel with net.mptcp.enabled=1: at
32761/32768 open files it stopped accepting connections — listen backlog
full, thousands of CLOSE-WAIT — while its own API kept answering over the two
keep-alive connections the monitor already held, and every local exporter
timed out for ten hours. The next day the leak came back on two hosts and the
profile, which only alerted, watched it happen again. The readiness and OTLP
watches now request a restart when the collector stops responding. FD usage
alone raises an alert unless the operator explicitly permits a restart.
| Watch | Signal | Action | Variable (default) |
|---|---|---|---|
ready |
GET /-/ready on the API port, on a fresh connection each cycle |
restart after 2 minutes unreachable; also the verify: true check a restart must pass |
host (127.0.0.1), port (12345) |
otlp |
disabled by default; when enabled, a valid empty JSON POST /v1/logs on the OTLP/HTTP receiver must answer 200 |
restart after 2 minutes of failure | otlp_port (4318) |
restart-if-fds-high |
the worst process against its own soft open-files limit; the sensor Sermo injects into every service | alert after 3 minutes above the limit; restart requires explicit permission | fds_limit (80%), a service key rather than a variable |
metrics (GET /metrics) stays graph-only. The service policy bounds the
retries — cooldown: 15m, max_actions: 2 per hour, backoff to one hour — so a
receiver that stops answering because of a downstream failure (a Loki that
rejects every push) costs at most two restarts an hour while the cause is
fixed. Like every remediation, these rules only simulate under
dry_run: true: the host's defaults.dry_run or the service's own dry_run
must be false for the restart to run, otherwise the daemon records
would restart events and nothing else.
A restart must also be able to clear the previous incarnation. An OpenRC
supervisor restart that leaves the old Alloy alive keeps :4318 bound, and the
new process starts without its OTLP receiver (bind: address already in use,
which Alloy does not retry) while /-/ready still answers. The profile
therefore declares one exact processes: identity per init backend —
main (user: ${user}, default alloy, the packaged systemd unit) under
systemd and main-openrc (user: ${openrc_user}, default root, OpenRC's
command_user) under OpenRC — with stop_policy.force_kill: auto, so the stop
phase signals a residual Alloy it can verify and never a process it cannot
name. Override the variable of your backend when the host runs Alloy as another
user; an exact selector that never matches a live process blocks every
operation. Survivors reparented to PID 1 are strays, which a restart reports
but does not signal; reap.kill_only_if authorizes sermoctl reap alloy --apply for an Alloy binary owned by one of those two users only.
Enable otlp only on an Alloy instance with an OTLP/HTTP receiver, and
override the port if needed:
name: alloy
uses: alloy
dry_run: false
variables:
otlp_port: 4319
watches:
otlp:
enabled: trueThe probe sends {"resourceLogs":[]} with Content-Type: application/json.
It creates no log records but detects HTTP rejections, server errors and
timeouts. It cannot prove that nonempty exports reach a downstream destination.
Opening /v1/logs in a browser sends GET and normally returns 405; that does
not indicate a failed receiver. The check remains opt-in because an Alloy
instance without an OTLP/HTTP receiver must not enter a restart loop.
optional: true on this check makes a failed observation a warning for service
health; it does not suppress its rule. Once enabled, an absent or
unreachable receiver must still restart. Leave the watch disabled on instances
without that receiver.
Every catalog service whose processes discovery can attribute gets the fds
sensor Sermo injects: a service metric watching the process closest to its own
soft RLIMIT_NOFILE, and a rule, restart-if-fds-high, that alerts after
three minutes above fds_limit (80%). Restart is disabled by default;
restart_on_fds_high: true explicitly enables it. No profile
writes it; the profiles that used to ship an alert-if-fds-high watch, or an
absolute fds ceiling such as 50000 that never fires for a daemon whose
limit is 32768, now rely on the injected sensor. The percentage is measured
per process — the process closest to its own limit — because the limit is per
process and that is the one that will fail accept() with EMFILE; see
Metrics and
fds_limit.
Services whose control group holds workload they do not own get no sensor: the
ones that delegate part of their tree (docker, containerd, virtnetworkd,
glusterd, ssh) automatically, and libvirtd through an explicit
fds_limit: false, because a guest's helper near its own limit describes the
guest rather than the daemon. A host that runs a daemon with a deliberately
small limit, or one that lives near it by design, tunes or vetoes the sensor
per service without a catalog edit:
name: mariadb
uses: mariadb
fds_limit: 95% # or false to drop check and rule
restart_on_fds_high: false # alert only (the default; overrides host opt-in)commands declares named auxiliary commands. Sermo never runs them as generic
checks, but the reserved names are consumed by features:
health— run by thesermoctl apps/libs/serviceslistings and the WebUI application list to decide whether an installed application is healthy. It uses the samepreflight.<name>thencommands.<name>lookup asversion, but only checks the exit code. When present, it takes precedence overversionfor app health;versionremains display-only.version(andversion_short) — run by thesermoctl apps/libs/serviceslistings to report a service's version, and each cycle by theversion.on_changemonitor (see Service health conditions). That monitor compares the numericversion_short, and an optionalversion.on_change.level(major/minor/patch, defaultpatch) selects at whicha.b.cgranularity a change should alert. The monitor inherits the service'sdry_runflag, so non-console notification delivery through that monitor is suppressed while the service is in dry-run mode. Top-levelevent_notifymay still deliver its alarm events. When both exist,preflight.versiontakes precedence overcommands.version. They also declareversionandversion_shortvariables with empty defaults for expansion; linked apps expose them to services as${app_version}and${app_version_short}. Other command-derived values can be declared withexport:, whose default source is trimmed stdout and whose default value is empty.
Any other entry is informational only. A run can assert its outcome, the same
way a watch hook or command check does: expect_exit (default 0, or a list
such as [0, 1]) and optional expect_stdout/expect_stderr matchers — a
substring or an {op, value} comparison (== != > >= < <= contains =~).
Reserved commands may also set user (username or numeric UID) to execute the
argv as that OS user when Sermo has permission to switch users.
commands:
version:
user: www-data
command: ["apachectl", "-v"]
timeout: 5s
expect_exit: 0 # optional, default 0
expect_stdout: { op: "=~", value: "Apache/2" } # optional: match the output