- Checks
- Ports
- TCP connections (tcp_connections)
- SSH terminal idle (ssh_idle)
- Terminal sessions (terminal_sessions)
- HTTP
- Cert
- Database connection (mysql / mariadb)
- SQLite integrity (sqlite / sqlite3)
- Replication (replication)
- SQL query (sql)
- Running database statements (db_queries)
- MongoDB query (mongodb-query)
- InfluxDB query (influxdb-query)
- Size growth (size)
- WebSocket (websocket)
- Clock drift (clock)
- Default route (route)
- Firewall rules (firewall_rules)
- Failed init units (failed_units)
- An unreachable ceiling is not a ceiling
- inotify limits (inotify)
- Disk throughput (hdparm)
- Missing devices
- What a device that stopped answering still reports
- Hardware sensors
- Count
- Log matches (log)
- Metrics
- Rules
- Remediation policy
Checks are single-shot probes under checks (and preflight, which reuses
the same schema). A checks entry flagged verify: true also runs as the
post-operation start verification. Supported types:
The complete set of single-shot check types is defined centrally. Tests lock that list against the builder dispatch and configuration validation, and lock the connection-protocol list below against the conn registry, so advertised check types cannot drift from the code. (Multi-metric watch forms such as net/icmp/swap and file/process watches expand on these primitives.)
Connection-protocol checks (MySQL, PostgreSQL, Redis, Docker, libvirt, etc.) are registered by their protocol name with common aliases (e.g. mysql/mariadb, fpm/php-fpm).
| type | style | passes when |
|---|---|---|
tcp |
health | a TCP connection to host:port succeeds |
tcp_connections |
condition | the number of local ESTABLISHED TCP sockets on port satisfies count {op, value} |
ssh_idle |
condition | interactive SSH terminals idle for idle_for, or protected terminal sessions, satisfy their configured count predicate |
terminal_sessions |
condition | the configured user's tmux or screen sessions satisfy a count/attached/detached predicate |
ports |
health | a set of host ports satisfy an open/closed expectation (see Ports) |
http |
health | the response matches expect_status (and optional headers/body/JSON, see HTTP) |
command |
health | the command exits with expect_exit (default 0) and its output matches optional expect_stdout/expect_stderr; optional user runs it as a specific OS user; on_change alerts when its output changes (e.g. a version), array form only |
config |
health | a config-test command (apachectl configtest, nginx -t, …) passes, and (with on_change) the config path is unchanged (see Service health conditions) |
service |
health | the backend status equals expect (active/inactive/paused/failed/unknown) |
file_exists |
health | a foreign flag/lock file exists (never under <runtime>/locks) |
file |
health | a path exists and is a regular file |
lockfile |
health | one service-created regular lockfile candidate exists — gate with requires: [service]; it does not block operations |
binary |
health | a path exists and is executable |
pidfile |
health | a pidfile exists and references a running, non-zombie process that the service's exe selectors or init backend claim — gate with requires: [service] so a missing/stale pidfile is an error only while the service is active |
socket |
health | one Unix socket candidate exists — gate with requires: [service] for sockets created by the service |
libraries |
health | all DT_NEEDED shared libraries of the binary can be resolved with the binary's ELF class and machine, each library's own RUNPATH/$ORIGIN searched first (native debug/elf, no ldd) |
process |
health | a process matching exe/user is in state (running/zombie/absent); an absent reading that a replaced binary explains names it and becomes a verdictless state (the service reads restart_required, not failed) |
process_policy |
health | every process of a user account satisfies the allow/deny policy (alert-only; see configuration.md) |
db_queries |
health | no MySQL/MariaDB/PostgreSQL statement matches the configured duration or resource thresholds; incidents tracked per statement (watch-only; see Running database statements) |
metric |
condition | a sampled metric satisfies op value (see Metrics) |
count |
condition | the number of entries in a directory satisfies op value (see Count) |
log |
condition | the lines appended to a log file (or glob) that match regex within within satisfy count {op, value} (see Log matches) |
storage |
condition | a filesystem's space/inode predicates hold (*_pct accepts %; *_bytes requires K/M/G/T) |
load |
condition | a load-average threshold holds (load1/load5/load15, optional per_cpu) |
users |
condition | the count of logged-in users (from utmp) satisfies count {op, value} |
process_count |
condition | the number of processes (host-wide, or filtered by user/exe/exe_dir) satisfies count {op, value} |
hdparm |
condition | a disk's hdparm read throughput crosses a threshold (read/cached MB/s) (see Disk throughput) |
sensors |
condition | hwmon hardware sensors cross a threshold (temp °C / fan RPM / voltage V) (see Hardware sensors) |
smart |
condition | a drive's SMART health/attributes and identity (failed verdict, reallocated, pending_sectors, crc_errors, media_errors, wear, temperature) (see Hardware sensors) |
storcli |
health | MegaRAID controller, enclosure, virtual/physical drive, cache protection, SMART/error and temperature health from StorCLI (see Hardware sensors) |
ssacli |
health | HPE Smart Array controller, logical/physical drive, cache, battery/capacitor, SMART/media-error and temperature health from SSA CLI (see Hardware sensors) |
raid |
condition | a Linux md software-RAID array is degraded/recovering (degraded/recovering/arrays) (see Hardware sensors) |
lvm |
health | every LVM volume group and logical volume reports healthy (see Hardware sensors) |
stale_binary |
condition | no attributed process runs a binary replaced on disk since it started (service-injected; also available explicitly) |
strays |
condition | the init unit's control group holds no process outside the configured selectors (service-injected; also available explicitly) |
edac |
condition | ECC memory errors from EDAC (ce correctable / ue uncorrectable) (see Hardware sensors) |
memory |
condition | system RAM vs the kernel's MemAvailable (used_pct/available_pct/available_bytes) |
pressure |
condition | kernel PSI stall time for cpu/memory/io (some_*/full_* avg10/60/300) |
fds |
condition | system file descriptors vs fs.file-max (used_pct/free/allocated) |
pids |
condition | the kernel PID table vs the tighter of kernel.pid_max and kernel.threads-max (used_pct/free/count) |
diskio |
condition | a block device's per-cycle I/O rates (util_pct/read_bytes/write_bytes/await_ms), plus cumulative read/written totals as readings |
conntrack |
condition | the netfilter conntrack table vs its max (used_pct/free/count) |
firewall_rules |
health | nftables/iptables has at least min_rules loaded rules (see Firewall rules) |
failed_units |
condition | the count of init units in a failed state satisfies count {op, value} (default > 0; see Failed init units) |
inotify |
condition | the per-user inotify instance/watch limits satisfy their {op, value} predicates (see inotify limits) |
route |
health | an up default route exists, optionally egressing a given interface (see Default route) |
clock |
health | local wall-clock offset stays within max_offset, measured against the configured NTP servers or (source: chrony) the local chronyd |
net |
condition | one interface metric (metric: state|speed|errors|address) holds — single-metric form of the net watch |
icmp |
condition | one ping metric (metric: state|latency) against host, optionally bound to an interface |
swap |
condition | one swap metric (metric: usage|io) holds — single-metric form of the swap watch |
zombies |
condition | the count of zombie processes satisfies count {op, value} |
oom |
condition | the kernel OOM-kill count rose by delta {op, value} since last cycle |
cert |
health | a TLS certificate is expiring/invalid, or its algorithm/issuer changed (see Cert) |
mysql / mariadb |
health | a MySQL/MariaDB server answers: with no credentials it reads the handshake greeting (liveness + version); with a user/password it authenticates and reads the version in one query (see Database) |
mongodb / mongo |
health | a connection to a MongoDB server authenticates, pings and reports its version and replica-set role for expect/on_change (see Database) |
postgres / postgresql |
health | a connection to a PostgreSQL server authenticates and responds (see Database) |
redis / valkey |
health | a connection to a Redis/Valkey server authenticates and answers PING; exposes role, replication, persistence and memory from INFO for expect (see Database) |
memcached / memcache |
health | a memcached server answers stats; exposes version, connections, hits/misses, items, bytes and evictions for expect (see Database) |
imap |
health | an IMAP server greets OK (anonymous) and, with credentials, LOGIN succeeds (see Database) |
pop / pop3 |
health | a POP3 server greets +OK (anonymous) and, with credentials, USER/PASS succeeds (see Database) |
smtp |
health | an SMTP server greets 220 + EHLO (anonymous) and, with credentials, AUTH PLAIN succeeds (see Database) |
smtp_acceptance |
health | a destination MX accepts this host's SMTP envelope through RCPT TO; the probe never sends DATA (see Database) |
nntp / nntps |
health | an NNTP server greets 200/201 (anonymous) and, with credentials, AUTHINFO USER/PASS succeeds (see Database) |
ftp |
health | an FTP server greets 220 (anonymous) and, with credentials, USER/PASS login succeeds (see Database) |
ssh |
health | an SSH server completes key exchange (anonymous: host key + banner); with credentials, login succeeds; on_change alerts on host-key change (see Database) |
fpm / php-fpm |
health | a PHP-FPM pool answers a FastCGI /ping with pong; an optional status_path exposes pool metrics for expect (Unix socket or TCP, see Database) |
dns |
health | a DNS server answers a query (NOERROR/NXDOMAIN) for query (see Database) |
ntp |
health | an NTP server answers with a synchronized time (server mode, stratum 1–15); exposes leap, precision, root delay/dispersion and reference id for expect (see Database) |
chrony / chronyd |
health | the local chronyd answers on its command port; exposes tracking (stratum, offset, leap, skew, frequency) and source counts for expect — use this, not ntp, for a chrony client that serves no NTP (see Database) |
snmp |
health | an SNMP agent answers a system GET (v2c community or v3 user/password); exposes sys name/contact/location/uptime for expect; on_change alerts on device-identity change (see Database) |
tftp |
health | a TFTP server answers an RRQ with a valid packet (DATA or ERROR) (see Database) |
ldap |
health | an LDAP directory accepts an anonymous bind, or a simple bind with credentials (see Database) |
ajp |
health | an AJP13 connector (e.g. Tomcat's 8009) answers a CPing with CPong (see Database) |
ipp / cups |
health | an IPP server (CUPS/cupsd) answers an IPP request with a valid response (see Database) |
rsync / rsyncd |
health | an rsync daemon sends its @RSYNCD: greeting (see Database) |
dhcp / dhcpd |
health | a DHCP server answers a DHCPDISCOVER with a DHCPOFFER (see Database) |
dhclient / dhcp-client |
health | a local DHCP client has UDP/68 bound in /proc/net/udp (see Database) |
rspamd |
health | an rspamd worker answers GET /ping with pong (see Database) |
libvirt / libvirtd |
health | a libvirt daemon answers RPC; exposes VM counts (domains.active…), node capacity and a VM's state for expect/on_change (see Database) |
dbus |
health | a D-Bus daemon answers GetId; named objects support read-only peer, introspection and scalar-property probes without activation (see Database) |
avahi / avahi-daemon |
health | the Avahi daemon answers GetVersionString over its D-Bus API (see Database) |
syncthing |
health | a Syncthing instance answers /rest/noauth/health with {"status":"OK"} (see Database) |
docker |
health | the Docker Engine answers /info, exposing container counts (running/paused/stopped), images and a container's state/health for expect/on_change (see Database) |
unifi / unifi-controller / unifi-network |
health | a UniFi Network controller answers GET /status with meta.rc == "ok" on 8443 (see Database) |
influxdb / influx |
health | an InfluxDB server answers /health (or /ping) and reports its version on 8086 (see Database) |
prometheus / prom |
health | a Prometheus server answers /api/v1/status/buildinfo (or /-/healthy) on 9090 (see Database) |
cloudflared / cloudflare-tunnel |
health | a Cloudflare Tunnel daemon answers /metrics on 60123 with cloudflared_ metrics (see Database) |
clamd / clamav |
health | a ClamAV daemon answers VERSION with its engine version (see Database) |
spamd / spamassassin |
health | the SpamAssassin daemon answers PING with PONG (see Database) |
nut / ups / upsd |
health | NUT's upsd answers VER; a UPS exposes its variables (status, battery charge/runtime, load, voltages) for expect/on_change (see Database) |
smb / samba / cifs |
health | an SMB/CIFS server negotiates (and, with credentials, authenticates) (see Database) |
acpid |
health | the ACPI event daemon accepts a connection on its Unix socket (see Database) |
fail2ban |
health | fail2ban-server accepts a connection on its control socket (see Database) |
lvmpolld |
health | LVM's poll daemon answers a hello request with OK over its socket (see Database) |
rpcbind / portmap / portmapper |
health | the RPC portmapper answers an RPC NULL call (see Database) |
nfs / nfs-server / nfsd |
health | an NFS server answers an RPC NULL call on 2049 (see Database) |
mountd / rpc.mountd / nfs-mountd |
health | the NFS mount daemon answers an RPC NULL call to MOUNT (100005) (see Database) |
statd / rpc.statd / nsm / nfs-statd |
health | the NFS status monitor answers an RPC NULL call to NSM (100024) (see Database) |
nebula / nebula-vpn |
health | a Nebula mesh-VPN node answers an unknown-tunnel packet with a recv_error on 4242/udp (see Database) |
openvpn / ovpn |
health | an OpenVPN server answers a hard-reset-client with a hard-reset-server on 1194 (see Database) |
rdp / ms-wbt-server |
health | a Remote Desktop server answers the X.224 connection negotiation (see Database) |
guacd / guacamole |
health | the Guacamole proxy daemon answers a select with a Guacamole instruction (see Database) |
asterisk / ami |
health | an Asterisk PBX sends its AMI Asterisk Call Manager/<version> greeting (see Database) |
sieve / managesieve |
health | a ManageSieve server sends its capability greeting ending in OK (see Database) |
mqtt |
health | an MQTT broker accepts a CONNECT (CONNACK return code 0) (see Database) |
amqp / rabbitmq |
health | an AMQP 0-9-1 broker sends a valid Connection.Start greeting (see Database) |
kafka |
health | a Kafka broker/controller answers an unauthenticated ApiVersions request; exposes the listener role (broker/controller) and produce_api/vote_api flags for expect (see Database) |
varnish / varnishadm |
health | the Varnish management CLI answers with its banner/auth challenge (see Database) |
ceph / ceph-mon |
health | a Ceph monitor sends its messenger ceph v… banner (see Database) |
glusterfs / glusterd / gluster |
health | a GlusterFS node accepts TCP connections to glusterd on 24007 (see Database) |
gluster_cluster |
health | the local Gluster CLI verifies peer membership, volume/bricks/self-heal state and optional heal limits (see Gluster cluster) |
openvswitch / ovs / ovsdb / ovsdb-server |
health | ovsdb-server serves a readable Open_vSwitch database over OVSDB JSON-RPC (see Database) |
sqlite / sqlite3 |
health | a SQLite database file passes PRAGMA integrity_check (see SQLite) |
replication |
health | MySQL/MariaDB replication is healthy: both replica threads run, lag optionally bounded (see Replication) |
sql |
condition | a SQL query's scalar result compares (== != > >= < <= contains =~) against a value (see SQL query) |
mongodb-query |
condition | a MongoDB document count / aggregation / command result compares against a value (see MongoDB query) |
influxdb-query |
condition | an InfluxQL (1.x) or Flux (2.x) query's scalar result compares against a value (see InfluxDB query) |
size |
condition | a file/directory grows by at least grow_by within within (runaway growth) (see Size growth) |
websocket |
health | a WebSocket endpoint completes the RFC 6455 opening handshake (see WebSocket) |
The storage check also verifies the mount of its path — see
storage and mount units.
process checks and process condition leaves match real UID/GID values read
from /proc/<pid>/status. A configured user: or group: name is resolved
through engine.user_lookup; if the name cannot be resolved it fails closed and
matches no process. Numeric UID/GID values avoid host identity-service ambiguity.
The command check asserts the command's outcome: expect_exit (default 0,
or a list such as [0, 1]) and optional expect_stdout / expect_stderr
matchers — a plain string requires that substring, or an {op, value} mapping
compares the trimmed output (== != > >= < <= contains =~):
checks:
queue-drained:
type: command
user: appqueue # optional: username or numeric UID on the host
command: [/usr/local/bin/queue-depth]
expect_exit: 0
expect_stdout: { op: "<", value: 100 } # fewer than 100 items queued
expect_stderr: "" # nothing written to stderruser runs the command as that OS user (Linux only). Sermo still executes the
argv directly, never through a shell; the daemon/CLI process must have permission
to switch user (normally by running as root), and an unresolved user or unsupported
runner fails the check closed. Like runuser, the command gets that user's
HOME, USER and LOGNAME, so clients find their own ~/.my.cnf or
~/.pgpass.
With a unit:, the first numeric token of the command's stdout publishes as
the check's value series in that unit — exim -bpc becomes a queue-depth
graph with two lines of YAML. Output with no leading number — including nan
and inf, which are not readings — simply records no sample (a gap), never a
failure; the exit code and matchers stay the verdict. The same holds for the
value of a sql, mongodb-query or influxdb-query result.
The same expect_exit / expect_stdout / expect_stderr fields are available
on a watch hook (then.hook) to validate the hook command's result, but
then.hook does not use user. A single-check watch may also declare
then.recover_hook with the same shape: it runs once on the failed-to-ok
edge — when a firing watch stops firing — never while healthy and never per
healthy cycle, with the same environment as then.hook plus
SERMO_EVENT=recovered. Dry-run reports it instead of executing it, and panic
mode suppresses it. The stateful file/process watches fire per path/PID and
do not accept it.
For systemd jobs that fail and immediately restart, sampling current failed
units can miss the failure. A bounded journal query can retain that evidence
in a host watch (or the same check in a service's watches):
name: recent-job-failures
interval: 1m
check:
type: command
command:
- journalctl
- --since=-10min
- --no-pager
- --quiet
- --output=cat
- --lines=20
- --grep=Failed with result '(timeout|exit-code|signal)'
expect_exit: [0, 1] # journalctl returns 1 when no lines match
expect_stdout: { op: "==", value: "" }
expect_stderr: { op: "==", value: "" } # do not accept journal read errors
timeout: 5sAn empty result passes; matching failures remain visible for ten minutes.
Choose a narrow expression for the jobs of interest. Unlike a log check,
this queries retained journal history even on its first cycle and after reload.
It requires readable systemd journals and a journalctl version with --grep.
expect_* is a single pass/fail assertion. To grade an otherwise-passing
command's output into warning or error, add an analyze: block. It references reusable rule sets from catalog/patterns/
(category patterns; sermoctl patterns catalog lists them all,
sermoctl patterns those in use) and can add or silence rules per
check:
checks:
config:
type: command
command: ["/usr/bin/named-checkconf"]
analyze:
use: [common, named] # inherit catalog/patterns sets, in order
silence: [deprecated] # drop inherited rules by id
rules: # service-local rules, evaluated FIRST (precedence)
- { id: zone-ok, match: "(?i)loaded serial", severity: ok }A pattern set is a patterns document (under catalog/patterns/) with an ordered, id'd rule list:
name: common
rules:
- { id: backup-now, match: "BACK UP DATA NOW", severity: error }
- { id: deprecated, match: "(?i)deprecated", severity: warning }matchis a Go RE2 regex ((?i)for case-insensitive);severityisokor one of the severity levels (debug|info|warning|error|critical); optionalstreamisstdout|stderr|both(defaultboth).- Evaluation: the resolved rule list is the check's local
rulesfirst (so a serviceokwhitelist or stricter rule overrides an inherited one), then theusesets in order (minussilenced ids). Per output line the first matching rule wins (anokmatch whitelists that line); the check's severity is the maximum over all lines. - Result:
errororcritical→ the check fails as required;warning,infoordebug→ the check fails as optional (does not block start/restart/reload/resume or drive remediation by itself); no match → the check passes. When the check declares noseverity:of its own, the matched grade is the failure's severity. A declared severity grades the check's failures: it may lower an advisory match, but never raises one — adeprecatedwarning stays a warning underseverity: error. The matchedpattern_idand line are in the result data. - Precedence: exit-code →
expect_*→analyze. The analyzer only grades a command that already passed its exit-code andexpect_*checks.
A resolved service that has preflight.config automatically receives the
periodic advisory check checks.configuration (default interval 15m). A
failure makes an active service warning (an error when preflight.config
declares severity: error, as Apache's does), remains outside SLA, and is visible
in the Web UI and through sermoctl status; sermoctl preflight SERVICE
reruns the required form on demand. The required preflight still blocks an
operation. The config.on_change block below is optional and adds persistent
change detection plus notifications; it is not required for health visibility.
A service can enable three standard health monitors with two short
declarative blocks — version: and config: — that reuse the version
and config commands the catalog service already defines (commands.version and
preflight.config). Sermo synthesizes a per-service monitor (a watch, built once
so change detection persists) from each:
# catalog service (e.g. apache.yml) — already defines these, unchanged:
commands:
version: { command: [apachectl, -v] }
preflight:
config: { type: command, command: [apachectl, configtest] }
# service (services/apache.yml) — opt into the monitors:
uses: apache
version:
on_change: { notify: [ops-email] } # alert when the version changes
# on_change: { notify: [ops-email], level: minor } # …only on major/minor bumps
config:
on_change: { notify: [ops-email] } # alert when the config is invalid…
path: [/etc/apache2/apache2.conf] # …or (optional) when this file changes-
Version changed —
version.on_changeruns the catalog service's version command and alerts (notifying the listed notifiers) when its version changes — an unexpected upgrade/downgrade. Needscommands.version(orpreflight.version) in the catalog service. The comparison is on the numericversion_short(a.b.c), so noise in the version banner (build dates, suffixes) does not trigger it. An optionallevelchooses the significant granularity:major— onlyachanges fire (1.4.2 → 1.9.0is ignored;1.x → 2.xfires).minor—aorbchanges fire (patch releases ignored).patch(default) — anya.b.cchange fires.
When the version output carries no parseable number, the monitor falls back to comparing the raw line so a change is never missed.
-
Config invalid / changed notification —
config.on_changeruns the catalog service'spreflight.configtest and alerts when it fails (invalid config); with apathit also alerts when a config file changes. A custompreflight:on the service replaces the catalog service'spreflight.config, and the monitor then uses that command, including itsuserfield when present. -
State not errored — the existing
servicecheck covers this: it alerts when the unit is not in the expected state (failed/unknown) or the backend cannot be queried.checks: state: { type: service, expect: active }
on_change.notify follows the usual notify precedence (omit to inherit the global
notify default, or none to suppress). A service-level dry_run: true
suppresses non-console notification delivery for these service-owned monitors;
wall still delivers. Top-level event_notify can independently deliver
their alarm events. The underlying command (on_change) and config check
types can also be used as host watch documents when you want a hook or a
standalone command.
On a multi-homed host (several NICs) a network check can be pinned to leave
through a specific interface with the optional interface field. The value
may be an interface name, an IP the interface carries, or its MAC, and
may be a single value or a list:
checks:
gw-via-wan:
type: icmp
host: 8.8.8.8
metric: state
expect: up
interface: eth1 # by name
db-on-mgmt:
type: tcp
host: 10.0.0.5
port: 5432
interface: 192.168.1.2 # by IP it carries
api-by-mac:
type: http
url: https://10.0.0.9/health
interface: "00:11:22:33:44:55" # by MAC
gw-redundant:
type: icmp
host: 8.8.8.8
metric: state
expect: up
interface: [eth0, eth1] # two uplinks
interface_match: any # any (default) | all- Optional. Omit
interface(the default) and the probe uses normal routing across all interfaces, exactly as before — nothing changes unless you set it. - Value forms.
eth0(name),192.168.1.2(an address on the interface), or00:11:22:33:44:55(MAC) — all resolve to the same interface. A list pins the check to several interfaces. interface_match(only meaningful with a list):any(default) — the check passes if the probe succeeds through at least one interface (failover/ redundant-link monitoring);all— it passes only if the probe succeeds through every listed interface (verify each path independently). The per-interface outcome is in the result data underinterfaces.- Mechanism. For TCP/UDP — including HTTP/3's QUIC socket — it binds the
socket with
SO_BINDTODEVICE, forcing egress through that interface regardless of the routing table; foricmpit binds the probe to the interface's IPv4 (theping -I <addr>mechanism). Linux only, andSO_BINDTODEVICEneedsCAP_NET_RAW(root) — if the interface does not exist or the daemon lacks privilege the check fails rather than silently using the wrong link. - Where it applies.
tcp,ports,icmp,websocket, and every connection-protocol check that dials TCP/UDP — native probes and driver-backed probes with a custom dialer such asmysql,postgres,mongodb,ldap,libvirt,redis,smtp,smtp_acceptance,dns,ntp,chrony,nfs,dhcp,openvpn,nebula,tftp, …, plus HTTP-based protocol probes such asinfluxdb/prometheus/cloudflared/syncthing/unifi/rspamd/ipp— honors the full list +interface_match. The standalonehttpcheck honors a single interface (the first listed). A check that dials a Unix socket instead of a host (chronywithsocket:,docker,libvirt, …) has no egress link, so aninterfacepin does not apply to it.
Any check may declare interdependencies so it is skipped (not counted, no
alert, shown as skipped) on a cycle where it should not apply:
checks:
port:
type: tcp
host: 127.0.0.1
port: 3306
query:
type: command
command: ["/usr/bin/mysqladmin", "ping"]
requires: [port] # skip while `port` is failing
skip_when_changed: ["/etc/my.cnf", "/etc/pam.d/mysql"] # skip while these changedrequires: [check, …]— skip this check while any listed check failed this cycle. This avoids cascading alerts: if MySQL'sportis down, the deeperquerycheck is skipped rather than also reported as failing. A dependency's own result counts even when its gate skips it, so in a chain wheredeeprequiresqueryandqueryrequiresport,deepis skipped wheneverqueryfailed.skip_when_changed: [path, …]— skip this check while any listed file differs from its acknowledged baseline (e.g. a config file or library was just updated). The baseline is re-acknowledged after a successful (re)start, so the check resumes once the service is reconciled.
Both accept a single value or a list. Gates are evaluated after the cycle's
checks run, so the probe still executes but its result is suppressed; use a check's
interval or remove it to avoid running it at all.
To restart a service when a library, file or app version is updated (the
other half of the example — "if the pam library was updated, restart"), use a
remediation rule with a changed: condition (or
restart_on_change: {paths: […], libraries: […], apps: […]}):
rules:
restart-on-pam:
type: remediation
if: { changed: { library: pam } } # or { path: /lib64/security/pam_unix.so }
then:
actions:
- type: alert
message: "${service} will restart after library change: ${change.library}"
- type: restartreports: declares what a check's result means, which is what decides
availability, SLA and how the dashboard labels it. It does not change how the
probe runs, and rules are unaffected: active: / failed: keep reading the
check's raw outcome.
| mode | meaning | dashboard | SLA |
|---|---|---|---|
health |
OK means the target is available | ok / fail |
counted |
condition |
OK means the condition fired, so availability is inverted | ok / fail |
counted |
state |
OK means a sensed state is present; neither side is good or bad | active / inactive |
n/a |
value |
the check measures and passes no judgement | measured plus the reading |
n/a |
Omitted, the check type supplies the default: health-style types
(tcp, http, service, the connection protocols, …) default to health, the
threshold and metric types to condition. The type is only a default — the same
type can be a health assertion in one service and a sensor in another, so the
check has the last word.
The state and value modes record no availability, so their SLA column reads
n/a rather than an empty series: there is no uptime to accumulate, and any
windows recorded before the mode was declared are dropped instead of ageing out.
A required check that starts failing records a firing event, and a recovered
event when it passes again — without needing a rule. Service events used to come
only from rules, so a required check with nothing bound to it moved the service
to failed on the dashboard and wrote nothing at all: the outage was visible
only to someone already looking at that service.
The reporting is edge-triggered on the observed value, so a condition that
persists is recorded once rather than every cycle. Optional, skipped and
verdictless (state/value) checks never raise one, a result reused from the
cache because of a per-check interval is not a fresh observation, and a check
some rule already reads is left to that rule so one incident does not become two
events. These events go to the event log only — notifications stay driven by a
rule's notify.
They keep their own colours — informative for a live state, muted for an idle
one — so neither is mistaken for the ok/fail verdict.
A condition check still reads ok / fail, not active / inactive: unlike
a state sensor, its firing side really is a problem and has to look like one.
A check that computes a figure every cycle and compares it against a limit has a
figure worth plotting: the limit says when to look, and the graph says what led
there. Every level check therefore graphs the numbers it already publishes with
no configuration at all — storage its used and inode percentages, load its
three averages, diskio its utilisation, throughput and await, and so on: each
type's own section below names its series.
This applies to a host watch exactly as it does to a service check: the watch
expansion draws the same panel on the same window selector, reading
GET /api/watches/{name}/metrics?metric=NAME. A watch has one check, so the
metric name alone identifies the series.
Some of a check's numbers are states, not magnitudes: RAID's degraded and
recovering are 0 or 1, and a dead-letter file's size matters only as "over the
threshold or not". A line chart through a boolean draws slopes that never
happened, so these render as an availability-style band instead — the same
strip, panel and window selector a service's SLA timeline uses — green while the
OK predicate holds, affected while it does not. Severity grades the failing
colour: an error band reads red on the usual down-share scale, a warning
band caps at amber, because an array rebuilding is a thing to watch, not an
outage.
The daemon records one state sample per cycle (up/down under the check's own
key), and the dashboard reads it from /api/.../sla?metric=NAME. Known boolean
states come banded out of the box — raid ships degraded (error) and
recovering (warning) — and a file watch with a size: predicate derives its
band automatically as that predicate's negation, an absent path counting
whatever absent_ok says. So both of these need no configuration at all:
watches:
raid-md126:
check: { type: raid, array: md126 } # two bands, zero YAML
watch-dead-letter:
check:
type: file
paths: [/root/dead.letter]
size: { op: ">", value: 0 } # band = size <= 0, absent ok
absent_ok: trueThe bands: block adjusts or extends that, on any check or watch check block:
check:
type: raid
bands:
degraded: { severity: warning } # keep the default ok, demote to amber
recovering: false # drop a default band
check:
type: load
bands:
load1: { ok: { op: "<", value: 8 }, severity: warning } # line chart -> bandA graph metric converted to a band must declare its ok: predicate — no default
exists for an arbitrary metric — and its value series stops being recorded: a
state draws as a band or as a line, never both. Check names may not contain :;
the band series is keyed check:metric.
The availability set is deliberately narrow — a threshold firing is not downtime
— but the operator who wrote the threshold may know better: a clock offset
breach, a firewall with its rules gone, or an unmounted filesystem is downtime
for whatever depends on it. sla: true on a watch's check block records its
verdict as an availability series and gives it the Availability panel; sla: false silences a type that would record by default. The verdict gates still
apply: a reports: sensor or a severity: warning advisory never enters the
series.
watches:
watch-clock-drift:
check: { type: clock, max_offset: 100ms, sla: true }
storage-mnt-backup:
check: { type: storage, path: /mnt/backup, used_pct: { op: ">=", value: "90%" }, sla: true }Most check types declare their graphable metrics statically, but some cannot: a
sql check's unit depends on its query, so five sensors on one service can
report MiB, seconds and a bare count. unit: declares it per check, and the
check's scalar result is then recorded and graphed like any other metric:
alert-if-replication-slot-backlog:
check:
type: sql
engine: postgres
query: "SELECT …" # returns MiB
op: ">"
value: 1024
unit: MiB # graph the result, labelled MiBIt is independent of reports: — a condition check keeps its threshold verdict
and gets a graph — and it adds to whatever the type already publishes.
Use state for a check that exists to answer "is this happening right now?"
rather than to assert something must hold. The catalog's backup watch is the
case: it detects a running backup so a guard can block a restart, and a host
with no backup in progress is normal, not unhealthy.
watches:
backup:
check:
type: process
reports: state # active / inactive, no verdict, no SLA
exe_any: [/usr/bin/pg_dump, /usr/bin/pgbackrest]
user: postgres
state: running # what is sensed — unrelated to `reports`
rules:
block-restart-during-backup:
type: guard
blocks: [restart]
if:
active: { check: backup } # still reads the raw outcome
then:
action: block
message: backup is running; restart deniedWithout it, such a check is a health assertion that fails whenever the state is
absent — a permanent red fail and a 0% availability series for what is in fact
the normal condition.
severity: grades how serious a failure is. It never changes the verdict — the
check fails exactly when it failed before — only how loudly that failure is
reported and who hears it. There are five levels, in ascending order:
| value | meaning | dashboard | daemon log | aggregate health and SLA | actions |
|---|---|---|---|---|---|
debug |
diagnostic chatter you opt into | amber row, debug badge |
level=DEBUG |
excluded | run |
info |
news worth recording, no action needed | amber row, info badge |
level=INFO |
excluded | run |
warning |
a degradation worth seeing, not worth waking anyone | amber, state warning |
level=WARN |
excluded | run |
error (default) |
an outage to act on | red, state failed |
level=ERROR |
counted against the target | run |
critical |
an outage that needs someone now | red, critical badge |
level=ERROR |
counted against the target | run |
debug, info and warning are advisories. An advisory still fires:
its for: window, its then.hook, its then.notify and its notification
cadence are untouched. What changes is that it stops competing for attention
with a real outage: it does not turn the watch row red, does not raise the
daemon's error count, does not hold against the service in the aggregate
health badge, and records no SLA series.
The level travels into every notification: the subject is tagged
[sermo][warning], [sermo][critical] and so on (an error keeps the plain
[sermo] subject it always had), the message carries SERMO_SEVERITY, and
each notifier's min_severity decides
whether it receives the message at all. That is how noise is routed: a chat
channel at warning sees advisories and outages, an on-call phone at
critical only what needs someone now.
Nothing that gates an automatic action reads it. Rule guards (active:,
failed:), remediation and start verification keep reading the check's raw
outcome, so an advisory can never be mistaken for a healthy target.
It can be declared at three levels, and the narrowest one wins:
name: net-enp1s0
severity: warning # widest: the whole watch, and its default
check:
type: net
interface: enp1s0
severity: error # narrower: this check and every metric under it
metrics:
state: # inherits error — a link going down is an outage
expect: down
for: { cycles: 3 }
errors:
severity: warning # narrowest: only this metric is an advisory
delta: { op: ">", value: 100 }
for: { cycles: 3 }A service check declares it the same way, beside reports: and timeout:. A
service watch declares it on the entry or its check: block; when the watch
desugars into a check plus a rule, the entry's severity grades that check. An
alert or remediation rule may declare its own severity: too; without one, the
rule takes the gravest grade among the failing checks its failed:/active:
conditions read (outside not:), and error when none carries one. A guard
reports no incident and accepts no severity.
optional: true remains the older, narrower spelling: it also keeps a failure
out of the service's availability, and it emits no warn-level log, but only
inside a checks: section. It does not keep the service out of the amber
state — an optional check that fails still reads the service warning, which is
what makes a mistuned threshold visible instead of silently swallowed.
The measurements worth grading as advisories are the ones that degrade
rather than break: hdparm throughput, icmp latency, and a net interface's
error counters. A disk answering from standby is the clearest case —
hdparm -t wakes it and times its spin-up, so it honestly reports a fraction of
a MB/s for a disk that is perfectly healthy.
Some check types grade their own findings when nothing declares severity:,
because they mix a verdict with early-warning counters:
service— a service proven not to be running (expect: activeand a status other thanactiveorunknown) iscritical. Anunknownstatus proves nothing and keepserror.process— a daemon expectedrunningthat is absent or a zombie iscritical.certandhttpcertificate options — a certificate inside itsexpires_in_days(cert_expires_in_days) window is awarning; one that already expired, is not yet valid, or whose file is gone iscritical; a broken chain or an unexpected algorithm, issuer or certificate change iserror. The gravest problem decides.commandon_changeandconfigfile-change detection — a changed output or configuration file isinfo; a failing command or an invalid configuration keepserror.smart— a predicate that holds (reallocated,pending_sectors,media_errors,crc_errors,temperature,wear,power_on_hours) while the drive's own verdict is PASSED or unknown is awarning: the disk works, it is starting to go. So is a SCSI/SAS drive's warning-class exception (sense code0Bh, reported ashealth=WARNING): a background scan found a medium error and the drive remapped the sector, or a temperature was exceeded — a lost sector, not a predicted failure. The FAILED verdict (an ATA prefail attribute past its threshold, a SCSI failure prediction5Dh) and a drivesmartctlcannot read stayerror.swap— every finding is awarning: a full swap area or steady paging is memory pressure to plan around, not an outage; thememorycheck owns the alarm for a host that runs out.storcli/ssacli— error counters on members whose state is still OK (media, other and predictive error counts, a volume's unrecoverable media errors) and thetemperaturepredicate arewarning; any state finding — a degraded controller, cache, battery, volume or drive, an inaccessible or inconsistent volume, unfinished parity or rebuild work, a drive's own SMART alert — iserror, and outranks the advisories beside it.lvm— a configuredfree_pctthreshold alone is awarning, including 0 % free: the VG has little room to allocate or grow LVs, but its filesystems can still have free space. Missing, partial or suspended volumes and thin-pool capacity thresholds remainerror; a volume fault outranks low VG headroom.analyzerules grade acommandcheck's output match (see Grading output withanalyze:).
A declared severity: on the watch, check or metric always wins over that
grade: severity: warning keeps every finding an advisory, severity: error
makes the counters outages again. One exception keeps a space ladder from
hiding a lost filesystem: on a storage check that declares both
mounted: and a space threshold, the declared severity: grades the
thresholds, and a wrong or absent mount is at least error. A mount-only
check keeps its declaration. So is a hung one: statfs on a hard network mount
whose server is gone blocks in the kernel, so the check gives it its timeout
(the engine's default_timeout unless set) and then reports the path
unavailable — error beside space thresholds — and later cycles fail at once
instead of stacking another blocked call on the same mount. The grade travels with the result, so the
row, the event's severity, the daemon log level, the notification subject and
the SLA follow it.
One measurement often deserves several grades: a disk at 80 % is worth a look,
at 95 % it is an outage, at 99 % someone must act now. levels: grades the
same sample on stricter thresholds, inside the check entry:
check:
type: storage
path: /
used_pct: { op: ">=", value: "80%" } # fires, graded by severity:
severity: warning
levels:
error: { used_pct: { op: ">=", value: "95%" } }
critical: { used_pct: { op: ">=", value: "99%" } }
for: { cycles: 3 }- The base threshold decides whether the check fails, and fails at the resolved severity (the declaration, or the type's self-grade).
- Each key of
levels:is a severity; its value restates the type's own threshold keys with a stricter value, in the same grammar. A level repeats only the keys it tightens. Acountorlogthreshold is restated in its nested form,{ count: { op, value } }, even when the base uses a top-levelop/value. An equality (==,!=) cannot tighten a shared key and is ignored. - When a level's thresholds breach too, the failure is raised to that level: the highest breaching level wins. Levels read the sample the check already took — nothing probes twice — and never grade an unavailable, skipped or verdictless result.
- What "breach" means follows the check's style. For a condition check
(
storage,memory,load,log,count,metric, …) every predicate of the level holds. Forcommandexpect_stdout, whose assertion states the passing side, the level loosens the bound and breaches when that looser assertion fails too: base{ op: "<=", value: 200 }warns above 200 messages, levelerror: { expect_stdout: { op: "<=", value: 1000 } }escalates above 1000. Output that is not a number never escalates. A connection protocol'sexpect:works the same way per field: a level restates some of the base's ordered fields,error: { expect: { maxmemory_used_pct: { op: "<", value: 85 } } }, and breaches when one of them fails too; a field the base does not bound with an ordered comparison cannot be graded. Forcertandhttp, a level{ expires_in_days: 7 }breaches when fewer than 7 days remain, including an expired certificate.smartORs its predicates, as its base does. levelslives in the check entry: a host or service watch'scheck:, a servicechecks:entry, or — fornet,icmpandswap— each metric block, since those grade one metric at a time. A preflight check and an inline rule probe read only the verdict and reject it.levels: falsedrops a block inherited from the catalog;levels.error: falsedrops one tier. A partial override merges into the inherited tier.
A level that cannot escalate is ignored, not rejected, so a configuration
keeps loading: one that is not strictly stricter than the threshold below it
(typical when an operator raises a catalog base threshold above a catalog
level), one that compares in the opposite direction or in the other metric
form (percentage versus absolute), or one that is not above the check's own
severity (levels.error under severity: error). For oom, failed_units,
edac and raid with no declared threshold, the one they fire on (any kill,
failed unit, uncorrectable error or degraded array: > 0) is the threshold
below the first level. A whole block under a reports: other than the type's
default is ignored too: the verdict it would grade is inverted or gone, which
is what happens when an override turns a catalog check into a sensor.
sermoctl config validate prints each as WARN and still succeeds, and
sermod logs it at load and reload. Raise the level with the base, or drop it
with false. Malformed blocks — an unknown severity name, a key that is not a
threshold of the type, an unsupported type — are validation errors. An
override that changes a check's type: never inherits the base's levels:,
which restate the old type's thresholds: it keeps none, or only the ones it
declares.
Types that accept levels:, and the threshold keys a level restates:
| Threshold form | Types | Level keys |
|---|---|---|
| level predicates | storage, memory, load, pressure, fds, pids, conntrack, inotify, diskio, sensors, hdparm, users, tcp_connections, ssh_idle, terminal_sessions, process_count, edac, raid, smart |
the type's predicate fields |
| single count | zombies, failed_units, log |
count |
| counter delta | oom |
delta |
| count or growth | count |
count, or delta in growth mode |
| per metric | swap |
usage: used_pct / free_pct / free_bytes; io: delta |
| per metric | net |
errors: delta |
| per metric | icmp |
latency with threshold: threshold |
| metric threshold | metric |
op + value |
| output assertion | command |
expect_stdout ({op, value} with >, >=, <, <=) |
| field assertions | connection protocols (redis, mysql, …) |
expect (its ordered {op, value} fields) |
| expiry window | cert |
expires_in_days |
| expiry window | http |
cert_expires_in_days |
Latency probes (tcp, http latency, a connection protocol's
expect_latency), composite health verdicts (lvm, storcli, ssacli), the host process, process_policy
and db_queries watches, and state, speed, address and change metrics have no single ordered
threshold and reject levels:.
A graded watch or rule escalates its open episode only when a graver level has
held for the owner's own window — the same for: or within: that opened the
episode. With for: { cycles: 3 } a disk must read above 95 % for three
consecutive cycles before the warning becomes an error; one spike does not
page anyone.
Inside an episode the grade only rises. Each escalation is announced once — a
firing event at the new severity and a new then.notify message — and the
episode keeps the gravest level it reached even when the value eases off, so a
value oscillating around a threshold costs at most one message per level per
episode. Reminders (notify_interval, event_notify.repeat_interval,
emission: every_cycle) repeat the held level. When the episode ends, the
recovered event and the recover hook's SERMO_SEVERITY carry that
high-water mark, and the recovery notification goes out at the gravest level
that was actually delivered, so exactly the notifiers that heard the incident
hear that it is over. A remediation rule whose action a cooldown holds back
still announces an escalation through its alert. A service check's health edge
has no window, so it escalates on the cycle its grade rises; after a daemon
restart it resumes from the last recorded grade.
tcp_connections is a local, condition-style check: it counts IPv4 and IPv6
TCP sockets in ESTABLISHED state whose local port equals port. It reads
/proc/net/tcp and /proc/net/tcp6; it does not open a network connection.
An unreadable table makes the check unavailable rather than report a partial
count. The one exception is a host booted with ipv6.disable=1: the kernel then
has no /proc/net/tcp6 and no IPv6 sockets, so only the IPv4 table is counted.
checks:
ftp-control-connections:
type: tcp_connections
port: 21
count: { op: ">", value: 5 }
reports: state # expose active/inactive without changing SLAIts result is a transport connection count, not an authenticated-user count. For FTP it covers control connections only: passive data sockets, TLS session state and login identity are deliberately not inferred. The same check is useful for connection thresholds on SSH, HTTP and other TCP services. Configure it on the host that owns the listening port; a reused port cannot be attributed to one service.
If /proc cannot be read, the check is unavailable rather than reporting zero.
When a guard references it, Sermo denies the operation conservatively. Set an
appropriate interval when using it as a long-running watch; a guard always
runs its check again immediately before an operation.
ssh_idle is a Linux condition-style check for interactive SSH terminals.
It reads utmp, the terminal input-access time and the process ancestry on that
TTY. sshd_exe is required and must name the exact resolved sshd executable;
this keeps local pseudo-terminals from being mistaken for SSH sessions.
It accepts one absolute path or a list of absolute paths when distributions use
different sshd locations.
checks:
idle-ssh:
type: ssh_idle
idle_for: 30m
sshd_exe: /usr/sbin/sshd
count: { op: ">", value: 0 }
reports: state
protected_processes:
deploy-account: { user: deploy }
dba-account: { group: database }
codex: { exe: /usr/local/bin/codex, user: deploy }
mysql-backup: { exe: /usr/bin/mysqldump, user: backup }protected_processes is a named map: entries are ORed, while its supplied
exe, user and group fields are ANDed. exe is an exact resolved
/proc/<pid>/exe path; user and group compare real UID and primary GID
(names or numeric IDs). An entry with only user or group is valid because
the candidate processes are already limited to that session's TTY. It never
widens service-process discovery or authorizes a signal.
count is the number of unprotected SSH sessions at least idle_for old.
protected_count is the number excluded by protected_processes, and
oldest_idle_seconds is the maximum input-idle duration among unprotected SSH
sessions. A screen or tmux window is the multiplexer's terminal, not
another SSH session, so it never adds to count or oldest_idle_seconds; a
protected process running in a window still adds to protected_count, even
when the multiplexer session is detached. To make a guard retain a protected
account or job, configure a second ssh_idle check with
protected_count: { op: ">", value: 0 } and reference it from the guard. The check itself never closes an SSH session.
SFTP/scp without a terminal and port-forward-only connections are deliberately
outside this check; use tcp_connections for transport connections. Terminal
atime depends on the host's atime policy. If utmp, a terminal, process ancestry,
an executable needed by a protection filter, or owner resolution cannot be read,
the check is unavailable; a guard therefore denies the operation rather than
assuming that no session is protected.
This fail-closed check behavior is independent from the dashboard inventory.
The dashboard may show a source as partial, retaining exactly verified SSH
rows while exposing each unverifiable terminal as an unavailable issue. On
systemd, a remote issue with a live utmp leader can be closed only through
login1's independently revalidated session identity; Sermo never signals that
uncertain PID directly. Such a terminal is not counted as a local console session.
terminal_sessions lists the active sessions of one explicitly configured
user through the installed tmux or GNU screen client. It is a
condition-style check; use reports: state to make it an informational
active/inactive sensor. The result exposes total count, attached and
detached sessions. The Web UI shows the individual sessions in its top-level
Sessions panel, together with SSH sessions. An administrator may close one row;
the action re-lists that configured namespace and requires the same session
generation before invoking the multiplexer client's exact close command.
An explicitly socket-configured tmux server that is live but has zero sessions
appears as empty; an administrator may close that server manually. Sermo
re-lists it, refuses the action if a session appeared, then invokes tmux's
kill-server command. If that tmux version leaves a stale socket, Sermo
removes only the same socket inode it observed before the close and only after
the namespace has disappeared. A namespace that is no longer present is
omitted from the Sessions panel.
watches:
tmux-sessions:
check:
type: terminal_sessions
multiplexer: tmux # tmux | screen
binary: /usr/bin/tmux # absolute path
user: deploy # account whose session namespace is queried
count: { op: ">", value: 0 }
reports: state
screen-sessions:
check:
type: terminal_sessions
multiplexer: screen
binary: /usr/bin/screen
user: backup
detached: { op: ">", value: 0 }
reports: stateAll queries run as the configured account with a bounded argv-only command; the
check itself does not inspect process names, enumerate unconfigured users,
attach to a session, or send a signal. A normal no-server/no-socket reply is an empty,
available sample. A tmux check may set an absolute socket: to query a
non-default server socket; screen has no socket: option. Command,
permission, timeout and malformed-output failures are unavailable rather than
reported as zero sessions.
For informational terminal inventory attached to SSH, set severity: warning
so an unavailable multiplexer remains visible without declaring SSH unavailable.
A successful reports: state sample still has no effect on service health.
A ports check probes several TCP ports on a host at once and evaluates a
quantified open/closed expectation. It is health-style (OK == true means the
expectation holds), so a watch over it fires its hook when the expectation breaks.
checks:
web-ports:
type: ports
host: 10.0.0.5 # default 127.0.0.1
ports: "80,443,1024-4000" # comma-separated single ports and inclusive ranges
expect: open # per-port desired state: open | closed | any (default open)
match: all # quantifier: all (AND) | any (OR) | none (NOT) (default all)
on_change: false # also fail when any port flips open<->closed between cycles
connect_timeout: 1s # per-port dial timeout (default 1s)expect is each port's desired state and match the quantifier over the ports in
that state: all = every port (AND), any = at least one (OR), none
= no port (NOT). So expect: open, match: all passes when every port is open;
expect: closed, match: any passes when at least one is closed. A port is
open when it accepts a TCP connection within connect_timeout, else closed.
expect: any skips the state expectation entirely — combine it with
on_change: true to alert purely on state transitions (a port that was open
becoming closed, or vice versa). Result data exposes open, closed, total and
changed. Ports are de-duplicated; a scan is capped at 16384 ports and runs
concurrently, but a large range of filtered ports (no response) is bounded only
by connect_timeout, so prefer tight ranges and a short timeout.
Like cert, the on_change detection is stateful (it remembers the previous
states across cycles). It works in service checks and host watches while the same
check instance is alive; the baseline is reset when the service worker or watch
is rebuilt, for example after a config reload.
Beyond the status code, an http check can send a method, headers and a body
(raw or JSON) and assert the response:
checks:
api:
type: http
url: "https://api.example.com/v1/health"
method: POST # any HTTP verb (default GET) — see below
headers:
Authorization: "Bearer ${token}" # any request headers
json: # request body as JSON (sets Content-Type
probe: true # automatically; or use `body:` for raw text)
expect_status: 200 # code, class (2xx), list, or { op, value }
follow_redirects: true # optional; false evaluates a 3xx as-is
expect_body: { op: contains, value: "ready" } # body comparison (see below)
expect_latency: { op: "<", value: 800 } # optional: response time in ms
proxy: "http://user:pass@squid:3128" # optional: route the request through a proxy (Squid)
expect_json: # optional: response JSON must match (dotted paths)
status: ok # equality (scalar)
data.replicas: { op: ">=", value: 2 } # operator: >, >=, <, <=, ==, !=, contains, =~
data.message: { op: contains, value: "healthy" }
data.version: { op: "=~", value: "^v[0-9]+" } # regex (Go/RE2)It passes (health-style, OK == true) when the status matches and every
assertion holds. method accepts any standard HTTP verb — GET (default),
HEAD, POST, PUT, PATCH, DELETE, OPTIONS, TRACE, CONNECT —
written in any case (it is normalized to upper-case); an unknown verb is
rejected at config validation. A request body/json is sent for any method
that carries one (POST/PUT/PATCH/…). http3: true sends the request
over HTTP/3 (QUIC) instead of TCP — see below. proxy routes the
request through a forward proxy such as Squid
(http://[user:pass@]host:port; http, https, socks5 or socks5h schemes
— credentials, when present, go in the URL). This both monitors that the proxy
forwards correctly and that the target is reachable through it; for an
https:// target the proxy is used via CONNECT, and certificate inspection
(below) still applies to the target's certificate. Without proxy, the check
(like every HTTP-based protocol probe) connects directly: the HTTP_PROXY,
HTTPS_PROXY and NO_PROXY environment variables are ignored.
json: marshals the value and sets Content-Type: application/json (override
it via headers). headers: {Host: app.example.com} selects an HTTP virtual
host while the URL still determines the connection address and TLS hostname.
body: sends a raw string. The response is only read when
expect_body/expect_json is set (capped at 1 MiB). expect_json looks up
dotted paths into nested objects. A scalar value is equality (==); a
{op, value} mapping uses an operator. Both use the same comparison as
expect_body: numeric when both sides parse as numbers (>, <, ==, !=,
…), otherwise string equality, contains substring matching, or a regex with
=~.
A response-body read error (including a timeout or premature disconnect) makes the check unavailable; a matching partial body does not count as success.
By default the check follows HTTP redirects using Go's standard client policy.
Set follow_redirects: false when the redirect itself is the health signal, for
example a local HTTP listener that intentionally redirects every request to
HTTPS.
Response comparisons. expect_body and expect_latency use an {op, value}
mapping. expect_status accepts either a code/class/list form or the same
{op, value} mapping. Operators are == != > >= < <= (numeric, or string for
==/!=), contains (substring) and =~ (Go/RE2 regular expression) — the
same operators as the sql check:
expect_status: { op: "<", value: 500 }— compare the status code numerically (in addition to the code/class/list forms).expect_body: { op: "=~", value: "^OK" }— compare the trimmed response body: numeric when both sides parse as numbers (>,<, …), otherwise string equality,containssubstring matching, or a regex with=~.expect_latency: { op: "<", value: 800 }— fail when the response time in milliseconds does not satisfy the comparison.
Result data carries status and latency_ms for use in rules/hooks.
One connection per probe. Every http probe (and every HTTP-based protocol
probe) opens its own connection and closes it after the exchange; it never
reuses a pooled keep-alive connection. A probe exists to exercise the target's
accept path each cycle: a daemon that stopped accepting connections — listen
backlog full, file descriptors exhausted — keeps answering on the sockets it
already owns, so a pooled connection opened before the collapse would report it
healthy while every new client times out. The cost is one TCP (and TLS)
handshake per probe; for an expensive endpoint, space probes with interval:
rather than sharing a connection. The http3 (QUIC) client is the exception:
it keeps its session.
On an https:// URL the same check can also inspect the server certificate
presented on the request connection, so one check covers reachability and TLS
health. Add any of these optional keys (they reuse the cert check's logic):
checks:
api:
type: http
url: "https://api.example.com/v1/health"
expect_status: 200
cert_expires_in_days: 14 # warn this many days before expiry
cert_verify: true # verify chain + hostname (default true here)
cert_on_change: false # alert on any rotation (leaf fingerprint)
cert_on_issuer_change: false # alert when the issuer changes
cert_on_algorithm_change: false # alert when the signature algorithm changesCertificate inspection activates when any cert_* key is present, and
requires an https URL — setting one on an http:// URL is a configuration
error. A certificate problem (expired/not-yet-valid, inside the
cert_expires_in_days window, failing verification, or a change between cycles)
fails the http check, keeping its health-style semantics (OK == true
means healthy), the same polarity as the standalone cert check. When
an HTTPS request redirects to plain HTTP, certificate inspection fails because
the final response has no certificate to inspect. When
redirects change the hostname, certificate verification uses the final URL's
hostname. When
inspection runs, the result data carries the same certificate fields the cert
check exposes (issuer, subject, dns_names, not_after, days_left,
fingerprint, …). To read the certificate even when it is expired or otherwise
invalid, the request skips transport-level verification and verifies the chain
manually; cert_verify: false disables that verification. The change conditions
are stateful (they remember the previous cycle). They work in service checks
and host watches while the same check instance is alive, and reset when the
service worker or watch is rebuilt. For raw TLS endpoints or local certificate
files, use the standalone cert check.
HTTP/3 (QUIC). Set http3: true to send the request over HTTP/3 (QUIC,
UDP) instead of TCP:
checks:
api-h3:
type: http
url: "https://api.example.com/health" # https only (QUIC is always TLS 1.3)
http3: true
interface: eth1 # bind the QUIC UDP socket
expect_status: 200
expect_latency: { op: "<", value: 300 }All the assertions above (status, body, JSON, latency, methods, and certificate
inspection) work the same over HTTP/3. The QUIC transport never falls back to
TCP, so a server that does not speak HTTP/3 — or a blocked UDP/443 — makes the
request fail and fires the check's alert/hook, which is how you monitor that
HTTP/3 stays available. The negotiated protocol is reported in result data as
protocol (e.g. HTTP/3.0; for normal checks it is HTTP/2.0 or HTTP/1.1).
HTTP/3 requires an https URL and cannot be combined with proxy (both rejected
at config validation). It can be combined with interface; the standalone
http check uses the first listed interface for both the request and certificate
inspection, and fails rather than falling back to the default route if the UDP
socket cannot be bound. Uses github.com/quic-go/quic-go (pure Go).
A cert check inspects TLS material — either a live TLS endpoint (host) or
a local file (path). It is health-style: OK == true means the certificate
or key material is acceptable, and any configured certificate problem makes the
check fail (OK == false). In rules, alert on certificate problems with
failed: {check: api-cert}. As a watch, the hook/notify fires when the check
fails.
checks:
api-cert: # live endpoint
type: cert
host: api.example.com # host XOR path (exactly one required)
port: 443 # optional, default 443
server_name: api.example.com # optional SNI + hostname to verify (default = host)
expires_in_days: 14 # optional: warn this many days before expiry
cert_verify: true # optional, default true: chain + hostname + validity
on_algorithm_change: true # optional: alert when the signature algorithm changes
on_issuer_change: true # optional: alert when the issuer (CA) changes / re-issue
on_change: false # optional: alert on any certificate rotation (fingerprint)
tls-keypair: # local file
type: cert
path: /etc/ssl/private/api.key # host XOR path
on_change: true # optional: alert if the file's fingerprint changes
rules:
alert-api-cert:
if:
failed: { check: api-cert }
then:
action: alert
message: "api.example.com certificate is invalid, expiring soon or changed"Host source. It fails when the certificate is expired or not yet valid,
expires within expires_in_days, fails chain/hostname verification
(cert_verify, on by default — catches self-signed, wrong host, expired chains), or —
between cycles — its signature algorithm, issuer or fingerprint changes.
A network/TLS error fetching the cert is not a certificate verdict: a
failed: rule on the check does not fire, so a certificate alert never doubles
as a reachability alert (use a tcp/http check for that). The cycle is still an
unavailable observation, like any probe that could not look: it counts
against the service's health and SLA, a watch reports it through its error
("check unavailable") event, and a guard reading it denies the action. Grade the
check severity: warning (see Severity) to keep an
unreachable endpoint amber instead of red.
File source (path). Reads and parses a local file, recognising natively (no
external tools): PEM certificate, certificate request (CSR), PKCS#1 / EC /
PKCS#8 private keys, PKIX public key, OpenSSH private key, and OpenSSH
public key (authorized_keys line). Certificates are checked for expiry/validity as
above; material that does not expire (keys, CSRs) fails only on
on_change/on_algorithm_change. A missing, unreadable or unparseable file makes
the check fail (a local configuration problem, unlike a transient network
error). cert_verify, port and server_name do not apply to files.
Result data exposes kind (certificate / certificate_request / private_key /
public_key / openssh_private_key / openssh_public_key / …), source,
signature_algorithm, public_key_algorithm, key_bits, subject and
fingerprint. Certificates additionally expose days_left, not_before,
not_after, issuer, serial_number (hex) and dns_names (SANs).
The change conditions are stateful (they remember the previous value across cycles). They work in service checks and host watches while the same check instance is alive; the baseline is reset when the service worker or watch is rebuilt, for example after a config reload.
Each check has an optional timeout (else engine.default_timeout) and an
optional interval to run it less often than the worker cycle — every
round(interval / resolution) cycles, reusing its last result in between (see
per-check interval).
A health check (tcp/http/service/command/cert/…) may also set
verify: true to double as the post-operation start verification: after a
successful start/restart/reload/resume the engine runs every verify: true
check up to five times, one second apart, within the operation timeout. It
fails the operation (postflight_failed) if a required one is still not OK —
its for/within window and any remediation are ignored, only the direct probe
result counts, and optional: true makes a failure a warning. This replaces the
retired postflight: section, so the health probe is defined once and serves
both periodic monitoring and start verification. verify: true is rejected on
condition checks (metric/storage/load/fds/…) whose OK does not confirm a
successful start.
A connection-protocol check connects to a server over its wire protocol and verifies it responds; the check type is the protocol name. A few conventions keep the per-protocol entries short:
tls(where listed) acceptsfalse(plaintext, the default),true(verified TLS) orskip-verify(TLS without certificate verification). Entries add only protocol-specific notes — the implicit-TLS port (e.g. IMAPS 993) or extra modes. The PostgreSQL sslmodes are accepted only bypostgres(and asqlcheck with a postgres engine); validation rejects them elsewhere, sotls: disablecan never switch TLS on for another protocol.- Auth is noted per entry; many protocols are anonymous.
socket(a Unix socket path) dials the socket instead ofhost/portfor the protocols that can reach a Unix endpoint:amqp,asterisk,avahi,chrony,clamd,dbus,docker,fpm,ftp,imap,kafka,libvirt,memcached,mqtt,mysql,nntp,nut,openvswitch,pop,redis,rsync,sieve,smtp,spamd,varnishand the socket-only daemons (acpid,fail2ban,lvmpolld). Validation rejectssocketfor any other protocol, which only dialshost/port.queryis the per-protocol lookup target (e.g. the DNS name fordns).- Shared text-protocol banner and line readers accept at most 64 KiB per line, including its terminator. Oversized lines fail the exchange.
HTTP-based protocol exchanges reject response-body read errors, including timeouts and premature disconnects; a matching partial response is not accepted.
Protocols, in the order of the table above:
-
mysql(aliasmariadb) — default port 3306;tlssupported.useris optional: with no user/password it reads the server's initial handshake packet (sent before auth) to prove liveness and report the version — no credentials, like the smtp/amqp greeting probes. With a user/password it authenticates and readsSELECT VERSION()in one round trip viagithub.com/go-sql-driver/mysql(the deeper check). An ERR handshake (host blocked, too many connections) fails the probe. -
mongodb(aliasmongo) — default port 27017;tlssupported.useris optional (MongoDB may run without auth); with credentials it authenticates againstauth_source(defaults todatabase, thenadmin). It connects directly to the configured node — replica-set discovery is off, so a secondary answers as itself and no primary is needed — verifies aping, and reads the version viabuildInfo. Ahello(with the legacyisMasteras fallback) reports the replica-setrole(primary/secondary/arbiter/standalone),set_nameandread_only, so anexpect:rule can assert e.g.role == primary. To run a query and compare a result, see the MongoDB query check. Usesgo.mongodb.org/mongo-driver. -
postgres(aliaspostgresql) — default port 5432;tlssupported, plus the PostgreSQL sslmodes (disable/require/prefer/verify-ca/verify-full).tls: trueissslmode=verify-full(certificate chain and host name checked against the system roots);skip-verifyissslmode=require(encrypted, unverified). Usesgithub.com/jackc/pgx/v5. -
redis(aliasvalkey) — default port 6379;tlssupported.useris optional (legacyrequirepassuses a password only, or no auth at all); a password-only check sendsAUTH <password>. VerifiesPING→PONGover RESP (no driver). A singleINFO all(one round trip, understood by every Redis, Valkey and KeyDB) then reports the serverversion(pair withon_version_change) plus health fields exposed forexpect::role,master_link_status(replicas),rdb_last_bgsave_status,aof_last_write_status,loading,used_memory,maxmemory,maxmemory_policy,evicted_keys,sync_full(full resynchronizations served to replicas),mem_fragmentation_ratio,connected_clientsanduptime_seconds. Three fields are derived:maxmemory_used_pctisused_memoryas a percentage ofmaxmemory(0when no limit is configured),keysis the key count summed over every database, andrejected_callssums the commands refused before they ran (anOOMundernoeviction,LOADING, a wrong arity) over every command of the reply'scommandstatssection; a server that predates the counter leaves it out.evicted_keys,sync_fullandrejected_callscount since the server started, so bound their rise withmax_increaserather than their value. Undermaxmemory-policy noevictiona full server rejects writes while still answeringPING, so assertmaxmemory_used_pctwhere that matters; pairkeyswithmax_increaseto catch a store filling up fast. -
memcached(aliasmemcache) — default port 11211;socketsupported (Unix socket),tlssupported. No auth (the ASCII text protocol). Sends a singlestatscommand and verifies the server answersSTATlines terminated byEND— proof the daemon is up. Reports the serverversion(pair withon_version_change) plus counters exposed forexpect::uptime,curr_connections,total_connections,rejected_connections,cmd_get,cmd_set,get_hits,get_misses,curr_items,total_items,bytes,evictions,limit_maxbytesandthreads(all numeric, so>/</==work). -
imap— default port 143;tlssupported (implicit TLS / IMAPS — use port 993).useris optional: with no credentials it verifies the server greets* OK; with a user/password it performs an IMAPLOGIN. RFC 3501. -
pop(aliaspop3) — default port 110;tlssupported (POP3S — use port 995).useris optional: anonymous verifies the+OKgreeting; with a user/password it performsUSER/PASS. RFC 1939. -
smtp— default port 25;tlssupported (SMTPS — use port 465; submission 587).useris optional: anonymous checks the220greeting +EHLO; with a user/password it performsAUTH PLAIN. RFC 5321. -
smtp_acceptance— default port 25. Resolves the MX records ofrecipient, connects through the selectedinterface, sendsEHLO, upgrades with STARTTLS, then testsMAIL FROMandRCPT TO. On acceptance it sendsRSETandQUIT; it has no code path that sendsDATA, so it never transfers or queues a message.helo,mail_from, andrecipientare required bare identities; the recipient must be a canary mailbox controlled by the operator.starttlsisrequiredby default or may beopportunistic. Result data includesmx_host,recipient_domain,starttls,accepted,smtp_stage, andsmtp_status(accepted,temporary,permanent,policy, orprotocol); an SMTP rejection also includessmtp_code, the boundedsmtp_reply, and the enhanced status when supplied. MX targets are tried in preference order, up to three: a transport failure advances to the next target, while the first SMTP or local-policy verdict is retained so a lower-priority MX cannot mask it. A STARTTLS certificate that fails verification counts as such a verdict (policyat thestarttlsstage), not as a transport failure. A null MX is reported as apolicyfailure. A pre-DATA acceptance detects connection, TLS and early reputation/policy blocks; it cannot prove content acceptance, inbox placement, DKIM signing or spam classification. -
nntp(aliasnntps) — default port 119;tlssupported (NNTPS — use port 563).useris optional: anonymous checks the greeting (200posting allowed /201prohibited — reported asposting_allowed); with a user/password it performsAUTHINFO USER/PASS. RFC 3977/4643. -
ftp— default port 21;tlssupported (FTPS — use port 990).useris optional: anonymous checks the220greeting; with a user/password it performsUSER/PASS(a password with no user logs in asanonymous). RFC 959. -
ssh— default port 22 (notls: SSH has its own transport crypto).useris optional: anonymous completes the key exchange to capture the server's host key (authentication then fails, which is expected); with a user/password login must succeed. Result data:fingerprint(SHA256 of the host key),host_key_algo,server_version,protocol. Seton_change: trueto alert when the host-key fingerprint changes — a possible re-key or man-in-the-middle. Usesgolang.org/x/crypto/ssh. -
fpm(aliasphp-fpm) — PHP-FPM over FastCGI. Setsocketto the pool's Unix socket (e.g./run/php/php8.2-fpm.sock), or usehost/port(default 9000) for a TCP pool. No auth. Performs a FastCGI request to/pingand expectspong, so the pool must haveping.path = /pingenabled. Each FastCGI response is limited to 1 MiB, including record headers and padding. Setstatus_path(the pool'spm.status_path) to additionally fetch the status page and expose pool metrics forexpect::pool,process_manager,active_processes,idle_processes,total_processes,listen_queue,max_listen_queue,max_active_processes,max_children_reached,slow_requests,accepted_connanduptime_seconds. -
dns— default port 53 (UDP). No auth. Sends anAquery forquery(defaultlocalhost) and verifies the answer:NOERROR/NXDOMAINpass (the server is up and speaking DNS);SERVFAIL,REFUSED, a timeout or a transport error fail. Result data: thercode, answer count and the resolvedaddresses(the answer's A/AAAA records, sorted and comma-joined) — soexpectcan require an actual resolution (rcode: NOERROR,answers: {op: ">", value: 0}) or a specific address (addresses: {op: "=~", value: "93\\.184\\..*"}). Setqueryto a name the server should answer (e.g. a zone it is authoritative for). Withresolvconf: true(instead ofhost, mutually exclusive) the probe asks the firstnameserverof/etc/resolv.conf— the server the system would ask first; with pppd'susepeerdns, the provider's resolver, which is how thepppdcatalog service verifies resolution through the uplink. If that resolver is local to the host (loopback such as127.0.0.0/8/::1, or any address assigned to a local interface), aninterfacepin is ignored for the DNS packet because the resolver must be reached locally. RFC 1035. -
ntp— default port 123 (UDP). No auth. Sends a client request and verifies the server answers in server mode with a synchronized stratum (1–15); a kiss-o'-death (stratum 0) or unsynchronized (stratum 16) reply fails. Result data:stratum, the clockoffset_seconds, theleapindicator (none/add-second/del-second/unsynchronized),precision_seconds,root_delay_ms,root_dispersion_msand thereference_id(a stratum-1 refclock label such asGPS, or the upstream server's IP). So anexpect:rule can assert e.g.leap == noneor aroot_dispersion_msceiling. RFC 5905. -
chrony(aliaschronyd) — default port 323 (UDP, chronyd's command port), orsocketfor its command socket (usually/run/chrony/chronyd.sock). No auth. This is the probe for a host running chrony: a chrony client-only configuration serves no NTP on port 123 at all, so thentpprobe cannot see it, while chrony's own command protocol reports the daemon's view of the clock. It reads the daemon's tracking state and its source counters. The probe never issues a command that changes anything; the one mutating command Sermo speaks is reachable only from a clock watch'sthen.makestepaction, never from a check, and only over the command socket. Result data shares thentpnames where the meaning matches —stratum,offset_seconds(chronyd's correction to the system clock),leap,root_delay_ms,root_dispersion_ms,reference_id— plussynchronized,reference_address,reference_time,reference_age_seconds,skew_ppm(already an unsigned error bound), plusfrequency_ppmandresidual_frequency_ppm— each of those two also asfrequency_abs_ppm/residual_frequency_abs_ppm, sinceexpect:has no absolute-value operator —rms_offset_seconds,last_offset_seconds,update_interval_secondsandsources,sources_online,sources_offline,sources_burst_online,sources_burst_offline,sources_unresolved. A daemon that is running but not yet disciplining the clock answers normally withsynchronized: false— that is a live daemon, so give sync loss its own watch with afor:window rather than expecting the probe to fail. UDP is the default because it is chronyd's monitoring-only channel: it refuses privileged commands, while the command socket is the fully privileged one.reference_idis chrony's hash of the peer address rather than the address itself, so it is reported as hex (an ASCII refclock label for stratum 1, as withntp); the real peer isreference_address— which is also the key to match on when a rule must work against both probes, since chrony derives an IPv6 peer's identifier by hashing it but uses an IPv4 peer's address verbatim, so the same upstream can read192.168.1.10fromntpandC0A8010Afromchrony. chronyd reports no version over this protocol, soon_version_changeis inert — use thechronydapp's--versionpreflight. Prefertype: clockwithsource: chronywhen you want drift thresholds and a graph rather than field assertions (see Clock drift). -
snmp— default port 161 (UDP). With nouserit uses SNMPv2c with a community string (password, defaultpublic— the anonymous/shared-secret model). With auserit uses SNMPv3 USM: apasswordadds SHA authentication (authNoPriv), otherwise noAuthNoPriv. It reads the system group; result data carriessys_object_id,snmp_version, the description (as the version banner) and — when the agent exposes them —sys_name,sys_contact,sys_locationandsys_uptime_seconds(assertable viaexpect:). Seton_change: trueto alert whensysObjectID(the device identity — model/firmware) changes. Usesgithub.com/gosnmp/gosnmp. -
tftp— default port 69 (UDP). No auth. Sends a read request (RRQ) forquery(defaultsermo-tftp-check) and verifies a valid TFTP packet: aDATAreply (the file is served) or anERRORreply (e.g. file not found) both pass. Only datagrams from the server's address count, and a transfer the probe opened is ended at once with a TFTPERROR, so the server does not retransmit. Result data: the reply kind and, for an error, the TFTP error code/message. RFC 1350. -
ldap— default port 389;tlssupported (implicit TLS / LDAPS — use port 636).useris optional: with no credentials it does an anonymous bind (a successful bind, or an LDAP-level rejection, both prove the directory is up — only a transport error fails); with a user/password it does a simple bind whereuseris the bind DN and must succeed. Result data: the bind mode and result. Usesgithub.com/go-ldap/ldap/v3. -
ajp— default port 8009 (TCP). No auth. Sends an AJP13 CPing and expects a CPong — the same liveness probe Apache/nginx use against Tomcat's AJP connector. -
ipp(aliascups) — default port 631;tlssupported (IPPS). No auth. POSTs an IPPCUPS-Get-Defaultrequest (asking only forprinter-name, so the reply stays small) over HTTP and verifies a valid IPP response — any parseable reply proves cupsd is up and speaking IPP. Result data: the IPP version and status. Encoding and parsing usegithub.com/OpenPrinting/goipp; Sermo retains HTTP transport, interface binding, TLS and response limits. RFC 8010/8011. -
rsync(aliasrsyncd) — default port 873 (TCP). No auth. Reads the rsync daemon's@RSYNCD: <version>greeting; receiving it proves the daemon is up. Asocketselects a Unix endpoint;tlssupports a TLS-wrapped daemon endpoint. Result data carries the protocol version. -
dhcp(aliasdhcpd) — default port 67 (UDP). Linux only. No auth. Sends aDHCPDISCOVERand verifies the server replies with aDHCPOFFER— proof it is up and handing out leases. It never sends aDHCPREQUEST, so no real lease is consumed. Two modes: setinterfaceto broadcast the DISCOVER out that link and discover any server (255.255.255.255); omit it to unicast tohost(a known server or relay). The client hardware address is a random, anonymous locally-administered MAC by default; setmacto use a fixed address (e.g. a server that only answers reserved clients). Result data: the offered IP, server id, subnet mask and lease time. Requires elevated privileges to bind the DHCP client port 68, like theicmpcheck; the host should not run a competing DHCP client on that interface. RFC 2131.Unlike the other probes, the per-interface mode pins the link with
IP_PKTINFOper datagram instead of binding the socket to the device, and filters replies by the link they arrived on. That is what lets it check a DHCP server running on this same host: such a server answers with a broadcast the kernel loops back, and the looped-back copy does not carry the LAN deviceSO_BINDTODEVICEmatches on, so a device-bound socket never sees the offer and the probe times out while the server answers every cycle. Egress is pinned just as strictly either way. Note the loopback caveat for the unicast mode: a server on this host will not answer a request aimed at127.0.0.1, or at the host's own LAN address, because the packet reaches it overlo, where it has no address pool — use the per-interface mode for a local server.checks: dhcp-broadcast: type: dhcp interface: eth0 # broadcast on this link (discovers any server) mac: "02:00:00:ab:cd:ef" # optional; default is a random anonymous MAC dhcp-unicast: type: dhcp host: 10.0.0.1 # unicast to a known server/relay (no interface)
-
dhclient(aliasdhcp-client) — default port 68 (UDP). Linux only. This is a local DHCP client check:dhclientreceives offers on UDP/68 and does not provide a request/response server protocol. The check reads/proc/net/udpand passes when it finds a local UDP socket bound exactly tohost:port. Unlike the other protocols,hostdefaults to the wildcard0.0.0.0(where a DHCP client binds), not127.0.0.1. It does not send packets and does not consume a lease. Setlease_file(the packaged catalog service defaults to/var/lib/dhcp/dhclient.leases; override it when your distribution stores ISC dhclient leases elsewhere) to also require an unexpired lease. Ifinterfaceis set, the lease must belong to that interface. -
rspamd— default port 11334 (the controller worker);tlssupported (HTTPS). No auth. SendsGET /pingand expects200with apongbody — the unauthenticated liveness endpoint every rspamd worker exposes (pointportat 11333 for the normal scanning worker or 11332 for the proxy). Result data: the rspamd version, read from theServerheader. -
libvirt(aliaslibvirtd) — opens an RPC connection to a libvirt daemon and reads its version; both succeeding prove libvirtd is up. It runs no write operation. Transport: with nosocket,hostorportit dials the local Unix socket/run/libvirt/libvirt-sock; setsocketfor a different path such as/run/libvirt/virtqemud-sockon modular libvirt hosts, or sethostand/orportto use plain TCP (default127.0.0.1:16509). TLS/SASL is not supported. Connect URI:queryselects the driver, defaultqemu:///system(e.g.lxc:///,xen://). No auth — local socket access is governed by the socket's permissions/polkit. Usesgithub.com/digitalocean/go-libvirt.Beyond liveness it exposes variables for conditions (best-effort — a driver that rejects them still reports up):
domains.active(running VMs),domains.inactive,domains(total), and node capacitynode.cpus,node.memory_mb. Setdomainto a VM name to also read itsdomain.state(running/paused/shutoff/crashed/…) anddomain.running;on_changethen alerts on that VM's state transitions, and an unknown domain fails the check. Result data also carries the libvirt version, connect URI, transport and hostname.checks: libvirt-local: type: libvirt # dials /run/libvirt/libvirt-sock expect: domains.active: { op: ">=", value: 3 } # alert if fewer than 3 VMs are running libvirt-modular: type: libvirt socket: /run/libvirt/virtqemud-sock query: "qemu:///system" libvirt-tcp: type: libvirt host: 10.0.0.4 # plain TCP on 16509 query: "qemu:///system" # optional connect URI (default qemu:///system) db-vm: type: libvirt domain: db01 # watch a single VM on_change: true # alert on its state transitions expect: domain.state: { op: "==", value: running }
-
dbus— connects to a D-Bus daemon and completes its SASL auth +org.freedesktop.DBus.Hellohandshake — which alone proves the bus is up — then callsorg.freedesktop.DBus.GetIdto read the bus UUID. With bothbus_nameandobject_path, it also resolves the well-known name withGetNameOwnerand probes the returned unique owner.probeselects one of three constrained read-only operations:peer(the default) callsorg.freedesktop.DBus.Peer.Ping;introspectparsesorg.freedesktop.DBus.Introspectable.Introspectand, whendbus_interfaceis set, requires that interface;propertycallsorg.freedesktop.DBus.Properties.Getand requires bothdbus_interfaceandproperty. Property values must be scalar (string, boolean, number or object path); useexpect.property_valueto assert the observed value. Every call setsNO_AUTO_START, so monitoring never activates a service and a call cannot race onto a replacement owner. By default, a name with no current owner is checked againstorg.freedesktop.DBus.ListActivatableNamesbefore it counts as a failure: an activatable name is installed and starts on demand —systemd-networkdroutes traffic perfectly whileorg.freedesktop.network1sits unactivated — so the check passes and reportsactivatable. Setrequire_owner: truefor a resident daemon:GetNameOwnermust then return a unique owner, and a merely activatable name fails. This detects a unit that remains active after losing its D-Bus connection. A name that is neither owned nor activatable really is absent, and still fails.require_owneris a boolean and requires the samebus_nameplusobject_pathtarget. There is no arbitrary-method mode. Omitting all target fields keeps the bus-only probe;bus_nameandobject_pathmust otherwise be set together. It runs no write operation. Target: defaults to the system bus (unix:path=/run/dbus/system_bus_socket); setsocketfor a different socket path, orqueryfor a full D-Bus address (unix:abstract=…,tcp:host=…,port=…). Socket-based, so there is no TCP port. No auth — access is governed by the socket's permissions. Result data: the bus id, address and the connection's unique name; a named-service probe also reportsbus_name,object_path,probe, its uniqueowner(theon_changefingerprint), and any configureddbus_interface,propertyandproperty_value. A name that passed because it is activatable reportsactivatableinstead of anowner. Usesgithub.com/godbus/dbus/v5. Withinterfaceand a TCPquery, use exactlytcp:host=…,port=…; address alternatives,nonce-tcpand other D-Bus TCP transport options are rejected so Sermo never silently drops egress-interface binding. Unix-only alternatives remain local and therefore need no interface binding.checks: dbus-system: # dials unix:path=/run/dbus/system_bus_socket type: dbus dbus-custom: type: dbus socket: /run/dbus/system_bus_socket # or use `query` for a full address login1: type: dbus bus_name: org.freedesktop.login1 object_path: /org/freedesktop/login1 require_owner: true udisks2: type: dbus bus_name: org.freedesktop.UDisks2 object_path: /org/freedesktop/UDisks2/Manager require_owner: true libvirt-dbus: type: dbus bus_name: org.libvirt object_path: /org/libvirt require_owner: true gdm-version: type: dbus bus_name: org.gnome.DisplayManager object_path: /org/gnome/DisplayManager/Manager require_owner: true probe: property dbus_interface: org.gnome.DisplayManager.Manager property: Version expect: property_value: { op: "=~", value: "^[0-9]+\\." }
-
avahi(aliasavahi-daemon) — the Avahi mDNS/DNS-SD (zeroconf) daemon, probed over its D-Bus API (org.freedesktop.Avahi). Connects to the system bus (SASL auth + Hello), resolves Avahi's unique owner without activation, and callsorg.freedesktop.Avahi.Server.GetVersionStringon that owner withNO_AUTO_START— a reply proves avahi-daemon is up and registered on the bus — reporting theversion(pair withon_version_change) and, best-effort, thehostnameand serverstate(runningwhen AVAHI_SERVER_RUNNING). Target: likedbus, defaults to the system bus; setsocketfor a different bus socket orqueryfor a full D-Bus address. Socket-based, no TCP port, no auth. Usesgithub.com/godbus/dbus/v5. The sameinterface+ TCPqueryrestriction asdbusapplies. -
syncthing— default port 8384;tlssupported (skip-verifycovers Syncthing's default self-signed GUI certificate). SendsGET /rest/noauth/healthand expects200with{"status":"OK"}— the unauthenticated liveness endpoint. With an API key inpassword(sent asX-API-Key) it also reads/rest/system/versionand reports the Syncthing version (os/archtoo); a rejected key fails the check. No user.checks: syncthing: type: syncthing host: 127.0.0.1 # tls: skip-verify # if the GUI is on HTTPS # password: "${env:ST_KEY}" # optional API key -> also reports version
-
unifi(aliasesunifi-controller,unifi-network) — a UniFi Network controller (Ubiquiti). Default port 8443, HTTPS-only with a self-signed certificate, sotlshere selects only verification: it is skipped by default; settls: trueto require a valid certificate. No user. SendsGET /status(the unauthenticated liveness endpoint) and expects200with JSONmeta.rc == "ok", reportingserver_version(pair withon_version_change) anduuid. Targets the self-hosted UniFi Network application; on a UniFi OS console (UDM/Cloud Key) the controller is proxied under/proxy/network/, which this check does not follow. -
influxdb(aliasinflux) — an InfluxDB server. Default port 8086;tlssupported (true/skip-verify→ https; plain HTTP by default). No auth. GETs/health(InfluxDB 2.x / 1.8+) and verifies a JSONstatusofpass, reporting the serverversion(pair withon_version_change); on older servers without/healthit falls back to/ping, which answers204with the version in theX-Influxdb-Versionheader. A liveness/version check; to run an InfluxQL query and compare a result, see the InfluxDB query check. -
prometheus(aliasprom) — a Prometheus server. Default port 9090;tlssupported (https). GETs/api/v1/status/buildinfoand requires HTTP 200 with a JSONsuccessstatus, reporting the serverversion(pair withon_version_change); on older servers it falls back to/-/healthy(liveness only). An optionaluser/passwordis sent as HTTP Basic auth (for a reverse proxy fronting the API). -
cloudflared(aliascloudflare-tunnel) — Cloudflare Tunnel's local metrics endpoint. Default port 60123;tlssupported (https, plaintext by default). GETs/metrics, requires HTTP 200, and verifies that the Prometheus text containscloudflared_metric names. This confirms the cloudflared daemon's own endpoint is responding instead of only checking that TCP accepts connections. -
clamd(aliasclamav) — default port 3310 (TCP), or a Unix socket viasocket(e.g./run/clamav/clamd.ctl). No auth, no TLS. Sends the clamdVERSIONcommand and verifies aClamAV <version>/…reply. Result data: the engineversion(the daily signature-database part is dropped, soon_version_changestays quiet across routine DB updates) and the fullversion_string. -
spamd(aliasspamassassin) — default port 783 (TCP), or a Unix socket viasocket. No auth. Sends a SPAMC/SPAMDPINGand verifies spamd answersSPAMD/<v> 0 PONG. Result data: the SPAMD protocol version. -
nut(aliasesups,upsd) — NUT (Network UPS Tools) upsd; default port 3493 (TCP),tlssupported (implicit TLS — upsd'sSTARTTLSupgrade is not used).user/passwordare optional: anonymously it sendsVERand reports the upsdversion(pair withon_version_change). With credentials itLOGIN-s to the UPS to verify access (USERNAME/PASSWORDalone are not checked by upsd).Set
upsto the device name (or omit it when the server has a single UPS — it is auto-detected) to read its variables into the result, where you alert on them withexpector on state changes withon_change. Exposed variables (when present):ups.status(the power/battery state —OLonline,OBon battery,LBlow battery,RBreplace battery,CHRG/DISCHRG…),ups.load,ups.temperature,ups.power/ups.realpower,battery.charge,battery.charge.low,battery.runtime/battery.runtime.low,battery.voltage,input.voltage,input.frequency,output.voltage,ups.mfr,ups.model. An unknownupsfails the check.checks: ups: type: nut host: 192.168.1.10 ups: myups # omit to auto-detect a single UPS user: monuser # optional (verifies access via LOGIN) password: ${env:NUT_PASS} on_change: true # alert on any ups.status transition expect: ups.status: { op: "=~", value: "OL" } # alert when not online (use =~: status is "OL CHRG") battery.charge: { op: ">", value: 30 } # alert when charge drops to 30%
on_changetracksups.status; for the upsd software version useon_version_change. Becauseups.statusis a space-separated flag list (e.g.OL CHRG), match it with=~rather than==. -
docker— the Docker Engine API. By default it talks to the local Unix socket/run/docker.sock; sethostand/orport(default127.0.0.1, port 2375 / 2376 withtls) for a TCP daemon, orsocketfor a non-default path. Nouser. It GETs/info(proving the daemon is up), reports the engineversion(pair withon_version_change), and exposes counts:containers,containers.running,containers.paused,containers.stopped,images, andwarnings(number of daemon warnings). Setcontainer(name or id) to also read that container'scontainer.status(running/exited/restarting/…),container.health(healthy/unhealthy/starting/none),container.running,container.restartcountandcontainer.exitcode;on_changethen alerts on its state/health transitions. An unknown container fails the check.checks: docker: type: docker # local socket by default expect: containers.running: { op: ">=", value: 4 } # alert if fewer than 4 are up containers.stopped: { op: "==", value: 0 } # alert on any stopped container warnings: { op: "==", value: 0 } # alert on daemon warnings web-container: type: docker container: web # watch one container on_change: true # alert on status/health transitions expect: container.health: { op: "==", value: healthy } container.restartcount: { op: "<", value: 5 } # alert on a restart loop
Most interesting conditions:
containers.running(expected services up),containers.stopped(crashed/exited containers), per-containerstatus/healthandrestartcount. The Docker check is read-only. To let Sermo start, stop, restart or resume that same container through the safe operation engine, add a service-levelcontrol: { type: docker, container: ... }block. andrestartcount(restart loops),warnings, andon_version_changefor engine upgrades. -
smb(aliasessamba,cifs) — default port 445 (TCP).useris optional. It first runs an SMB2NEGOTIATE(proving the server is up) and reports the negotiated dialect as theversion(2.0.2/2.1/3.0/3.0.2/3.1.1— pair withon_version_change), theprotocolfamily (SMB2/SMB3) and whether signing is required. With auserit then authenticates over NTLM (a failed login fails the check), counts the shares (shares), and — if a share is named inquery— verifies it can be mounted (share_access). The domain may be embedded inuser(DOMAIN\useroruser@domain). The NEGOTIATE is native; the authenticated session usesgithub.com/cloudsoda/go-smb2.checks: fileserver: type: smb host: 10.0.0.9 user: "WORKGROUP\\monitor" # optional; enables NTLM auth + share checks password: "${env:SMB_PASS}" query: "data" # optional: verify this share mounts
-
acpid— the ACPI event daemon. Socket-only (no TCP port;hostandportare rejected; defaults to/run/acpid.socket, override withsocket). It is an event broadcaster with no request/response protocol, so the check is the connect itself: a successful connection proves acpid is listening (a stale socket left by a dead daemon refuses the connection). It reads nothing — reading would block until an ACPI event — and there is no version. No auth. -
fail2ban— fail2ban-server. Socket-only (hostandportare rejected; defaults to/run/fail2ban/fail2ban.sock, override withsocket). Its Python pickle command protocol is not worth reimplementing for a liveness check, so — likeacpid— the check is the connect itself; it exchanges no commands. No auth. -
lvmpolld— LVM's poll daemon. Socket-only (hostandportare rejected; defaults to/run/lvm/lvmpolld.socket, override withsocket). Unlike acpid/fail2ban it is probed by protocol: it speaks LVM's generic daemon framework, so the check sends ahellorequest and verifies the daemon repliesOK, also guarding against a different LVM daemon (lvmetad, dmeventd) by the reported protocol name. Result data: theprotocolandprotocol_version(the handshake exposes no lvm2 software version). No auth. -
rpcbind(aliasesportmap,portmapper) — default port 111 (UDP). No auth. Sends an ONC RPC NULL call (RFC 5531/1833) to the portmapper program (100000 v2) and verifies that the requested program is present: a successful reply or a program-version mismatch passes, while an unavailable program, denied reply or another acceptance status fails. Result data carries therpc_status. The same NULL-call probe backs thenfs/mountd/statdchecks below. -
nfs(aliasesnfs-server,nfsd) — an ONC RPC NULL to the NFS program (100003) over TCP (record marking), likerpcbind; default port 2049. A version-mismatch reply (e.g. an NFSv4-only server answering a v3 NULL) still passes. NFS-family RPC replies are limited to 1 MiB in total across fragments. -
mountd(aliasesrpc.mountd,nfs-mountd) — the NFS mount daemon: an ONC RPC NULL to the MOUNT program (100005) over TCP, likenfs. No fixed well-known port — mountd registers a (often random) port with rpcbind; default 20048, overrideport(find it withrpcinfo -p <host>). -
statd(aliasesrpc.statd,nsm,nfs-statd) — the NFS status-monitor (NSM, used for lock recovery): an ONC RPC NULL to the NSM program (100024), likemountd. Default port 662; same no-fixed-port caveat — overrideport(rpcinfo -p <host>). -
nebula(aliasnebula-vpn) — a Nebula mesh-VPN node. Default port 4242 (UDP). No auth. A real tunnel needs a CA-signed certificate, but a node answers a data packet for a tunnel index it does not know with a plaintext recv_error (telling the sender to re-handshake), so the check sends a Nebulamessagepacket carrying a random index and verifies the node replies with arecv_errorechoing it — proof the node is up, with no credentials. The reply is governed by the node'slisten.send_recv_errorsetting (defaultalways); a node set tonever— or toprivatewhen probed from a public address — stays silent and reads as down, so probe lighthouses/nodes from an address their config answers. -
openvpn(aliasovpn) — an OpenVPN server. Default port 1194;transportselects the transport (udp, the default, ortcp— match the server'sproto). No auth. The first step of the OpenVPN handshake is unauthenticated (TLS comes after): the check sends aP_CONTROL_HARD_RESET_CLIENT_V2carrying a random session id and verifies the server answers with aP_CONTROL_HARD_RESET_SERVER_V2acknowledging it. Result data: thetransport. Caveat: the reset only gets a reply from a server withouttls-auth/tls-crypt; those HMAC-wrap (or encrypt) control packets, so a bare reset is dropped — silence is then expected and is not proof it is down. -
rdp(aliasms-wbt-server) — default port 3389 (TCP). No auth. Sends an X.224 Connection Request with an RDP Negotiation Request and verifies the server answers with an X.224 Connection Confirm; a negotiation failure still counts as up (the server answered). Result data: the negotiatedsecurityprotocol (rdp= standard RDP security,tls,hybrid= CredSSP/NLA,hybrid-ex). MS-RDPBCGR; the negotiation precedes authentication, so no credentials. -
guacd(aliasguacamole) — default port 4822 (TCP). No auth. Opens the Guacamole handshake by sending aselectinstruction for a protocol (query, defaultvnc) and verifies guacd replies with a well-formed Guacamole instruction — anargsreply (protocol available) or anerror(e.g. plugin missing) both prove guacd is up. Result data: the selected protocol and the replyopcode. -
asterisk(aliasami) — default port 5038 (TCP);tlssupported (AMI over TLS). No auth. On connect, Asterisk's Manager Interface sends anAsterisk Call Manager/<version>greeting before any login; reading it yields the managerversion(result data also carries the fullbanner). Pair withon_version_changeto alert on an Asterisk upgrade. -
sieve(aliasmanagesieve) — default port 4190 (TCP);tlssupported (implicit TLS). No auth. On connect the server sends a greeting of capability lines terminated by anOKresponse (RFC 5804); reading it and seeing theOKproves the server is up. TheIMPLEMENTATIONcapability is reported as the serverversion(aNO/BYEgreeting, e.g. a connection-limit refusal, fails the check). -
mqtt— default port 1883 (TCP);tlssupported (MQTTS, port 8883). Performs an MQTT 3.1.1CONNECThandshake and verifies the broker answersCONNACKaccepting the connection (return code 0). With no credentials it is an anonymous connect;user/passwordauthenticate (MQTT 3.1.1 allows nopasswordwithout auser; the check fails before connecting). A refused CONNACK (e.g.not-authorized,bad-username-or-password) fails the check with the reason; result data: theconnackstatus. -
amqp(aliasrabbitmq) — default port 5672 (TCP); no auth. Sends the AMQP 0-9-1 protocol header and verifies the broker's unprompted Connection.Start method. Reports the brokerversionplus best-effortproduct,platformandcluster_namefields forexpect/on_version_change. -
kafka— default port 9092 (TCP);tlssupported. No auth. Sends anApiVersionsrequest (API key 18, v0), which a broker or a KRaft controller answers before authentication, and verifies the reply's correlation id matches — proof the peer speaks the Kafka wire protocol. From the advertised API set it derivesrole(brokerwhen the data-plane Produce API is present,controllerwhen the RaftVotequorum API is, and Produce is not) and theproduce_api/vote_api(yes/no) flags, plusapi_countanderror_code— all assertable viaexpect. Used by thekafka-broker(9092,expect role=broker) andkafka-controller(9093,expect role=controller) catalog services. -
varnish(aliasvarnishadm) — default port 6082 (TCP, the Varnish-Tmanagement CLI). No auth. On connect varnishd sends a CLI response (a<status> <length>line and a body); status 200 carries the banner (with the version) and 107 is an authentication challenge (a CLI secret is set) — either proves the management CLI is up. Any other status, or a body shorter than its declared length, fails the check. Result data: thecli_statusand, for a banner, the Varnishversion. The CLI secret authentication is not performed (liveness only). -
ceph(aliasceph-mon) — default port 3300 (TCP, the Ceph monitor's messenger v2; use port 6789 for the legacy v1). No auth. On connect a Ceph daemon sends a messenger banner (ceph v2\nfor v2,ceph v027for v1); reading aceph vbanner proves it is a Ceph endpoint. Result data: themessengerversion (v1/v2). The banner precedes the authenticated handshake, so no credentials. -
glusterfs(aliasesglusterd,gluster) — default port 24007 (TCP, the glusterd management daemon). No auth. The check only establishes a TCP connection: GlusterFS deliberately leaves the RPC NULL actor in its handshake program unimplemented, and calling it logs a false error in current glusterd releases. This is a node liveness check; usegluster_clusterfor the local node's view of cluster health. -
openvswitch(aliasesovs,ovsdb,ovsdb-server) — default port 6640 (TCP, the Open vSwitch configuration database serverovsdb-server), or a Unix socket viasocket(commonly/run/openvswitch/db.sock);tlssupported (SSL). No auth. Issues an OVSDB (RFC 7047)list_dbsJSON-RPC request, requires theOpen_vSwitchdatabase, then uses atransactselect to require its readable root row. Its optionalovs_versionis reported as theversionwhen populated; result data also carries thedatabaseslist. Theovsdb-servercatalog profile additionally runs the read-onlyovsdb-client needs-conversioncheck every 10 minutes. It passes only when the live database schema matches the installed/usr/share/openvswitch/vswitch.ovsschema(outputno);yesor a client error reports the schema watch unhealthy without converting the database. A separate dailydatabase-integritywatch runs the read-onlyovsdb-tool show-logover/var/lib/openvswitch/conf.db(override the profile'sdatabasevariable when the file lives elsewhere). It parses every log record and works with standalone, active-backup and clustered files; an unreadable, truncated or malformed log fails the watch. For a clustered deployment, an operator may additionally runovsdb-tool check-cluster DB...with copies from the different cluster members to check self-consistency and cross-consistency; Sermo does not infer or collect those remote files.
A sqlite check verifies a local SQLite database file is healthy by running
SQLite's integrity check. It is a local file check (not a network protocol).
checks:
app-db:
type: sqlite
path: /var/lib/app/app.db # required
quick: false # optional: true runs the faster PRAGMA quick_checkIt passes (health-style, OK == true) when PRAGMA integrity_check reports
ok. A missing/unreadable file, a file that is not a SQLite database, or
reported corruption fails the check with the detail. The file is opened
read-only, so the check never modifies it. quick: true runs
PRAGMA quick_check (faster, skips some per-row checks) for large databases.
A gluster_cluster check is a local, read-only view of the node's own
cluster state — not a network protocol, so it is documented here rather than in
the protocol list. It runs the installed gluster --mode=script --xml client
under the check timeout, using the host's existing Gluster management
credentials and trust configuration, and never changes cluster state. Use
glusterfs for a plain node-liveness probe of the management port.
Supply one or both of peers and volumes: every configured peer must be
present, connected and a cluster member, and a disconnected peer returned by
Gluster also fails the check. A disconnected peer is reported once under its
Gluster hostname, even when configured through an alias. Every configured
volume must exist, be started and
have exactly its configured number of bricks online. Only entries of
gluster volume status with a brick directory count as bricks; the quota,
bitrot, scrubber, snapshot and NFS daemons listed beside them do not, and a
brick the status omits (its peer dropped out) is reported as missing. self_heal: true requires
a running self-heal daemon; max_heal_entries and max_split_brain_entries are
optional non-negative limits (use 0 to require no pending entries).
watches:
cluster:
interval: 2m
check:
type: gluster_cluster
peers: [node-b, node-c]
volumes:
images:
bricks: 3
self_heal: true
max_heal_entries: 0
max_split_brain_entries: 0The heal limits only apply to replicated volumes: gluster volume heal fails on
a pure distribute volume, so declaring them for one makes the whole check report
Unavailable.
A brick that cannot answer a heal query — Gluster reports - instead of a count,
usually for a disconnected one — is listed as its own issue naming the brick and
the status Gluster gave for it. It never makes the check Unavailable: the peer,
volume, brick and self-heal state already collected stays reported.
A replication check watches MySQL/MariaDB replication — a master-master pair
above all — through the server's own status rows. It is a health check
(OK == true means every watched connection replicates): both replica threads
must be Yes, and the lag may carry an explicit behind bound. The status
query is tried newest-vocabulary-first (SHOW ALL SLAVES STATUS,
SHOW REPLICA STATUS, SHOW SLAVE STATUS), so one check covers MariaDB,
MySQL 8 and older servers, whichever column spelling they answer with.
watches:
db-replication:
category: database
interval: 1m
check:
type: replication
engine: mariadb # mysql | mariadb (default mariadb; same wire protocol)
host: 127.0.0.1 # same connection fields as the mysql checks,
# or socket: /run/mysqld/mysqld.sock
user: root
password: "${env:SERMO_MYSQL_PASSWORD}"
behind: { op: "<", value: 60 } # optional: fail when the lag breaks this bound
# connection: primary # optional: scope to one MariaDB multi-source connection
replication_control:
start: true # offer the manual START REPLICA repair in the dashboard- With no
connection, every replication connection the server reports must be healthy; the published lag is the worst across connections. Withconnection:, only that MariaDB multi-source connection (or MySQL channel) is judged, and an unknown name fails the check by name. - Thread states publish as
io_stopped/sql_stopped(0 ok / 1 stopped) and render as SLA-style bands; the lag graphs asbehind_seconds. A stopped thread's message and reading quote the server's ownLast_IO_Error/Last_SQL_Error— the text a DBA acts on. An IO thread inConnectingstate is not replicating and does not pass. replication_control.start: trueoffers the manual repair in the dashboard: Sermo revalidates live status, runsSTART REPLICA(or the engine's older spelling, or the MariaDB named-connection form) exactly as the manual process would, then re-reads status until both threads run. Withoutconnectionon a MariaDB server with named connections it runsSTART ALL SLAVES, since plainSTART SLAVEwould start only the default connection while the check covers them all. It is an explicitly requested admin action behind confirmation — never autonomous — and it cannot skip or discard replication events.- A server with no replication configured fails the check: declaring the watch asserts replication exists.
A sql check runs a query against a database and compares its scalar result
(the first column of the first row) against a value. It is condition-style
(OK == true means the comparison holds), so in rules active: {check: …}
fires on it. It uses the same connection fields as the MySQL/PostgreSQL checks
and opens SQLite databases read-only.
checks:
jobs-backlog:
type: sql
engine: postgres # mysql | mariadb | postgres | postgresql | sqlite | sqlite3
host: 127.0.0.1 # mysql/postgres: host/port/user/password/database/tls
user: monitor
password: "${env:PGPASS}"
database: app
query: "SELECT count(*) FROM jobs WHERE state = 'queued'"
op: ">" # == | != | > | >= | < | <= | contains | =~
value: "100"
schema-version:
type: sql
engine: sqlite
path: /var/lib/app/app.db # sqlite: a path, opened read-only
query: "SELECT value FROM meta WHERE key = 'schema'"
op: "=~" # regular expression (Go/RE2)
value: "^v[0-9]+$"- Operators:
>,>=,<,<=compare numerically (result andvaluemust parse as numbers);==/!=compare numerically when both are numbers, otherwise as strings (equal/different);=~matches the result againstvalueas a Go (RE2) regular expression. - Engines:
mysql/mariadbandpostgres/postgresqluse the same connection fields as their protocol checks (host/port/user/password/database/tls) and require auser;sqlite/sqlite3take apathand open it read-only. Amysql/mariadbengine also acceptssocket(a Unix socket path, e.g./run/mysqld/mysqld.sock) in place ofhost/port; thereplicationcheck honours it the same way. - Result data carries
engine,query,op,threshold, the rawresultstring and, when numeric, avaluefor hooks/rules. A query error, a missing database or aNULLresult fails the check. The check only reads — point it at a read-only user.
A db_queries watch lists the statements a MySQL, MariaDB or PostgreSQL server
is running and alerts per statement on duration, CPU or memory thresholds. It
is watch-only (a host watch under watches: or a service watch); it cannot be a
service checks: entry or a preflight. The catalog mysql, mariadb and
postgres services ship one as alert-if-query-long-running (see
services).
watches:
long-queries:
interval: 30s
check:
type: db_queries
engine: mariadb # mysql | mariadb | postgres | postgresql
socket: /run/mysqld/mysqld.sock # mysql/mariadb only; or host/port
defaults_file: /root/.my.cnf # mysql/mariadb only
min_duration: 5m # duration-only alert
exclude_users: [backup] # optional filters: which statements alert
severity: warning
summary: "query ${id} by ${user} on ${database} running ${elapsed}: ${query}"engine(required):mysql/mariadbreadinformation_schema.PROCESSLIST,postgres/postgresqlreadpg_stat_activity. A MySQL/MariaDB server's flavour and whether it answers the thread columns (MariaDBTID, MySQLperformance_schema.threads) are detected once, so a sample is one query; a failed sample detects them again.host,port,user,password,database,tls: the usual database connection fields; PostgreSQL requiresuser.socket(MySQL/MariaDB only): a Unix socket path instead ofhost/port.defaults_file(MySQL/MariaDB only): an option file read natively for the fields the check leaves unset (see Credentials below).min_duration: without resource thresholds, a statement running at least this long is an incident. With resource thresholds, this is an optional minimum age before resources are evaluated. If supplied, it must be positive. At leastmin_durationor one resource threshold is required. Only a duration-only watch marks a statement as long-running; passing a resource watch's age gate does not itself mark or alert the statement.cpu,cpu_thread,memory: optional{op, value}resource predicates, evaluated against each statement's own sample. Any matching resource predicate triggers the statement's incident, aftermin_durationwhen set and only within the account/database filters.cpuis a percentage of all host CPUs;cpu_threadis a percentage of one logical CPU (100% means one fully occupied CPU). CPU thresholds accept numbers or%in 0..100.memoryis bytes and requires a size suffix, for example256MiBor1G. It is MariaDB's per-connection memory or a PostgreSQL backend's resident memory; MySQL does not supply per-connection memory.users/exclude_users,databases/exclude_databases: only statements of (not of) these accounts, in (not in) these databases, alert.states(PostgreSQL only): thepg_stat_activity.statevalues listed (default[active]; addidle in transactionto see open transactions).max_query_length: statement text kept per row, in characters (default1000).max_rows: rows published to the dashboard, longest first (default50).timeout,severity,summary: as for any check.
For example, this service watch alerts when a statement at least 30 seconds old
exceeds either 90% of one CPU or 256 MiB. Add cpu to watch total CPU
share as another independent limit. On eight logical CPUs, a thread using 100%
has a total CPU share of 12.5%.
watches:
resource-heavy-queries:
interval: 30s
check:
type: db_queries
engine: mariadb
socket: /run/mysqld/mysqld.sock
defaults_file: /root/.my.cnf
min_duration: 30s
cpu_thread: { op: ">", value: "90%" }
memory: { op: ">", value: 256MiB }
exclude_users: [backup]
severity: warningThe same check works in a host watch. Omit min_duration to evaluate resource
limits as soon as the readings are ready; CPU still needs two samples. CPU
requires a local database and a readable, verified server thread. PostgreSQL
memory also requires a local backend; MariaDB reports its connection memory
through SQL, including for remote connections. Missing readings are unknown,
never zero: a known matching resource can fire, but otherwise a missing reading
makes the result unavailable. An open incident stays open through missing
samples. These watches do not support for:, within: or levels:.
If the statement's OS thread changes, CPU and IO need a new baseline on the
verified replacement thread; counters from different threads are never combined.
The dashboard's check predicates show every configured resource threshold and
the minimum age when present.
What is listed. The watch's own connection never appears. MySQL/MariaDB
skip idle and server threads: Sleep, Daemon, Binlog Dump*, Slave_*,
Connect and Register Slave commands, the system user and
event_scheduler accounts, and rows with no statement text. PostgreSQL lists
client backends only, in the configured states. The filters decide which
statements alert; the dashboard still lists every statement, so a filtered one
is visible but never an incident. MariaDB reports elapsed time to the
millisecond; MySQL's TIME counts whole seconds.
Credentials. defaults_file reads the [client], [client-server],
[client-mariadb] and [mysql] groups of a MySQL option file (user,
password, socket, host, port) for every field the check does not set
itself — the same credentials the mysql client run by root uses. A missing
file means no options, and user defaults to root. MySQL needs the PROCESS
privilege to list other accounts' threads; PostgreSQL needs pg_monitor (or
superuser) to read other roles' statement text.
Per-statement incidents. Each matching statement emits one firing event
at the watch's severity, whose
message names the engine, connection id, elapsed time, user, database, client
host, command/state and the statement text; a resource alert also names the
matching thresholds and their measured values. A resource incident recovers
when a valid sample no longer meets any configured threshold, and can fire
again if usage rises. When the statement finishes — or
its connection starts another one — the watch emits recovered. Each statement
is its own incident in event_notify
(the event's check is query:<key>), so two slow statements are two alerts,
each recovered on its own, and every target's min_severity applies. The
watch's published result is failing while any statement is over the threshold,
with count, long_count and oldest_seconds in its data. Resource watches
also publish matched_count, and unknown_count when readings are missing.
Bounded samples. Each cycle evaluates at most the 500 oldest statements.
The sampler reads one extra row to detect an incomplete list; exactly 500 rows
can still be complete. An incomplete sample publishes complete: false and
cannot report a healthy watch just because none of the sampled statements
matches. Counts describe the sampled rows. Open incidents for unlisted statements
stay pending until a later observation establishes their recovery. A different
statement observed on the same connection confirms that the previous one ended,
even in a partial sample. Pending incidents are persisted separately from the
current statement list and are not
shown as live sessions or considered for automatic cancellation. The Sessions
source reads partial when the list is incomplete or required resource readings
are missing, even when another statement has already crossed a threshold.
Restarts. The alerted statements are persisted with the watch's sample. A
restarted or reloaded sermod restores them as already announced: a statement
still running is not alerted twice. The initial observation-only cycle does not
fire hooks, send notifications or close incidents, including when a connection
has started a replacement statement. Pending recoveries survive another restart
and are dispatched on a live cycle once the observation confirms them.
The Sessions expansion follows the same statement across refreshes, including small changes in MySQL's estimated start time. Its stable presentation identity is separate from the latest identity used to re-verify cancellation; a new statement on the same connection starts collapsed.
Failures and stopped services. A connection or query failure publishes the
watch as unavailable and reports one availability incident; it recovers on the
first good sample. A service watch skips its sample while the service is not
active, so a stopped database is not reported as a broken probe.
Hooks and notifiers. then.notify and then.hook work as on any watch,
once per statement incident; when a notifier announced the statement live, its
recovery is dispatched too. The hook environment adds SERMO_DB_ENGINE,
SERMO_DB_QUERY_ID, SERMO_DB_USER, SERMO_DB_NAME, SERMO_DB_HOST,
SERMO_ELAPSED_SECONDS, SERMO_QUERY and SERMO_CHANGE (long for a duration
alert, threshold for a resource alert, cleared when resource usage recovers,
ended when the statement finishes). Measured samples also supply SERMO_CPU,
SERMO_CPU_THREAD (percentages) and SERMO_MEMORY (bytes). summary: may use
${id}, ${user}, ${database}, ${host}, ${query}, ${elapsed} and
${value} (elapsed duration), plus ${cpu}, ${cpu_thread} and ${memory}
(bytes); a missing resource renders as unavailable. levels: and
then.action are rejected: each statement uses the watch's own severity.
Redaction. Statement text is redacted before it is stored, shown, logged or
handed to a hook: the secret after IDENTIFIED BY, PASSWORD(...),
SET PASSWORD, MASTER_PASSWORD/SOURCE_PASSWORD and PostgreSQL's
PASSWORD '...' becomes '***'. Whitespace is collapsed to single spaces, then
the text is cut to max_query_length.
Killing a statement. The dashboard's Sessions panel and
sermoctl sessions kill cancel one listed statement of a service watch
(kill_query, a guard-blockable manual action; see
safety). A service watch may also opt
into an automatic kill:
# services/mariadb.yml
uses: mariadb
watches:
alert-if-query-long-running:
check:
min_duration: 5m
then:
kill_query:
after: 30m # required; at least check.min_duration
mode: query # query (default) | connection
users: [report] # users and/or databases: required
databases: [analytics]
policy:
cooldown: 10m # required, positive
max_actions: 3
max_actions_window: 1hThe kill is never shipped by the catalog and only exists on a service watch: it
runs through the service's operation engine. At most one statement is killed per
cycle, the longest eligible one first, and only one the watch counts as long —
past min_duration and within the check's own users/exclude_users/
databases/exclude_databases filters, so it never stops a statement it was
told to ignore — that has run at least after and matches the kill's
users/databases (both lists, when given, must match). Immediately before
acting, the engine re-reads only the target and requires the same statement
identity and the same after, selector and filter condition; otherwise it does
nothing. The watch's own policy: paces the kills, separately from the
service's restart budget, and only a kill that stopped a statement spends it (a
held-back kill is reported once per statement). dry_run: true reports would kill_query once per statement, panic mode holds the kill back until it clears,
and a guard that blocks kill_query blocks it. A PostgreSQL session idle in transaction runs no statement to cancel: mode: query refuses it, so watching
that state with an automatic kill needs mode: connection. policy: is only
accepted together with then.kill_query.
Resource thresholds cannot be combined with then.kill_query: its fresh target
read revalidates duration and selectors, but cannot revalidate a CPU rate
without a second sample. Use a separate duration-only watch for automatic
cancellation. Resource watches still permit manual cancellation through the
usual identity verification, operation locks and guards.
A mongodb-query check runs a MongoDB query, compares a scalar result with
value, and is condition-style (OK == true means the comparison holds).
It uses the same connection variables as the mongodb connection check
(host/port/user/password/database/tls, plus auth_source), the same
direct connection to that node, and the official MongoDB driver. Three query shapes are supported:
checks:
failed-jobs: # 1) document count
type: mongodb-query
host: 127.0.0.1 # host/port/user/password/database/tls/auth_source
user: monitor
password: "${env:MGPASS}"
database: app
collection: jobs
filter: '{"status":"failed"}' # optional JSON filter; default {} (count all)
op: "<" # == | != | > | >= | < | <= | contains | =~
value: "10"
queued-jobs: # 2) aggregation pipeline (scalar at `result`)
type: mongodb-query
database: app
collection: jobs
pipeline: '[{"$match":{"state":"queued"}},{"$count":"n"}]'
result: "n" # dotted path into the first result document
op: ">"
value: "100"
connections: # 3) database command (scalar at `result`)
type: mongodb-query
database: app # command runs here; defaults to admin
command: '{"serverStatus":1}'
result: "connections.current" # dotted path into the reply
op: "<"
value: "5000"- Query shapes (exactly one): a
collection(+ optional JSONfilter) compares the matching document count; acollection+ JSONpipelineruns an aggregation; acommandruns a database command.pipelineandcommandextract a scalar at the dottedresultpath (a collection count needs noresult).filter/pipeline/commandaccept relaxed extended JSON (so$oid,$date, etc. work). A collection query requires adatabase;commanddefaults toadmin. - Operators behave exactly as the
sqlcheck's (>>=<<=numeric;==/!=numeric-or-string;containssubstring;=~RE2 regexp). - Auth: with a
user, credentials are checked againstauth_source(defaultdatabase, thenadmin). The check only reads — point it at a read-only user. - Result data carries
mode,op,threshold, the rawresultand, when numeric, avaluefor hooks/rules.
An influxdb-query check runs an InfluxDB query, compares a scalar result
with value, and is condition-style (OK == true means the comparison
holds). It uses the influxdb connection variables (host/port/user/
password/tls). The language selects the query API:
influxql(default) — InfluxDB 1.xGET /queryagainst adatabase.flux— InfluxDB 2.xPOST /api/v2/queryagainst anorgwith atoken.
checks:
cpu-load: # InfluxQL (1.x)
type: influxdb-query
host: 127.0.0.1 # host/port/user/password/tls (https when tls set)
user: monitor # optional: sent as HTTP Basic auth
password: "${env:INFLUXPW}"
database: telegraf # required for influxql
query: "SELECT mean(usage_user) FROM cpu WHERE time > now() - 5m"
op: "<" # == | != | > | >= | < | <= | contains | =~
value: "80"
disk-flux: # Flux (2.x)
type: influxdb-query
language: flux
host: 127.0.0.1
tls: true # InfluxDB 2.x is usually https
org: my-org # required for flux
token: "${env:INFLUX_TOKEN}" # required for flux (Authorization: Token …)
query: >
from(bucket: "telegraf")
|> range(start: -5m)
|> filter(fn: (r) => r._measurement == "disk" and r._field == "used_percent")
|> mean()
op: "<"
value: "90"- Scalar selection. InfluxQL returns rows of
[time, …]; by default the result is the last column of the first row of the first series (the aggregate value, sincetimeis first). Flux returns annotated CSV; by default the result is the_valuecolumn of the first data row. Setcolumnto read a named column in either mode. A query that matches nothing fails the check ("no value"). - Operators behave exactly as the
sqlcheck's (>>=<<=numeric;==/!=numeric-or-string;containssubstring;=~RE2 regexp). - Auth. InfluxQL: a
user/passwordis sent as HTTP Basic auth; an optionaltoken(1.8+/2.x compatibility) is sent asAuthorization: Token …and takes precedence. Flux: thetokenis required. The check only reads — point it at a read-only user/token. - Result data carries
language,query,op,threshold, thedatabase/orgin use, the rawresultand, when numeric, avaluefor hooks/rules. A query error (e.g. unknown database, bad token) fails the check.
A size check watches a file or directory and alerts when it grows by at
least grow_by within the within window — useful to catch a runaway log, a
disk-filling spool or a leaking cache. Only increases trip it: a steady or
shrinking path passes. It is condition-style (OK == true means "grew too
fast", so active: {check: …} fires) and stateful. Growth history persists
while the service worker or watch is alive and resets when that worker/watch is
rebuilt, for example after a config reload.
watches:
log-runaway:
check:
type: size
path: /var/log/app.log # a file, or a directory (recursive sum of file sizes)
grow_by: 1GB # alert if it grows at least this much…
within: 1h # …within this sliding window
then:
notify: [ops-email] # and/or a hook(The examples in this file use compact global watches: maps. In a file under
paths.watches, write the same watch as a name: log-runaway document and keep
the inner fields at the top level.)
(Note within here is the size check's own field — the duration of its
growth window — not the watch-level within: {cycles|duration, min_matches}
firing window, which a size watch normally does not need.)
Each cycle it samples the path's size (a file's bytes, or the recursive sum of
regular-file sizes under a directory), keeps the samples seen in the last
within, and compares the current size against the oldest one still in the
window. It fails when current − baseline ≥ grow_by. The first cycle only
baselines (no alert). grow_by uses the same size grammar as every other size
field (free_bytes, expand.by): an explicit K/M/G/T suffix (optional
B/iB), binary units (1G = 2³⁰), with plain byte counts rejected. It must
be positive and fit a signed 64-bit byte count. Result
data carries current_bytes, baseline_bytes,
growth_bytes, the window and value (the growth) for hooks/rules. A
directory walk skips hidden descendants by default; set include_hidden: true to
include them. A hidden path named directly is always sampled. Point it at a bounded
path.
A websocket check verifies a WebSocket endpoint completes the RFC 6455 opening
handshake: it sends the HTTP Upgrade request and checks the server answers
101 Switching Protocols with a Sec-WebSocket-Accept matching the sent key
(so it confirms a real WebSocket server, not just any HTTP 101).
checks:
realtime:
type: websocket
url: "wss://example.com/socket" # ws:// | wss:// | http:// | https://
# tls: skip-verify # accept a self-signed cert (wss/https)
# origin: "https://example.com" # optional Origin header
# subprotocol: "chat" # optional Sec-WebSocket-Protocol
# headers: { Authorization: "Bearer ${token}" } # optional extra headersIt passes (health-style, OK == true) when the handshake completes. ws/http
connect in plaintext; wss/https use TLS (tls: skip-verify accepts a
self-signed certificate). The default port follows the scheme (80 / 443) unless
the URL gives one. Result data carries the negotiated subprotocol. Probed
natively (no external library).
checks:
db:
type: mysql # any protocol from the supported-types table above
# Auth depends on protocol: postgres requires user; mysql/mongodb can probe without it;
# redis/imap/pop/smtp may be anonymous; fpm/dns/amqp use no auth.
host: 127.0.0.1 # default 127.0.0.1
port: 3306 # default: the protocol's port (mysql 3306, postgres 5432)
user: monitor # optional/required by protocol
password: "${env:DB_PASS}" # resolved from the environment at load (never store secrets in plaintext)
database: "" # optional
tls: false # optional (see per-protocol values above)
timeout: 5s # optional (engine.default_timeout)It passes (health-style, OK == true) when it connects, authenticates as
user, and the server answers a ping. Result data exposes protocol, host,
port and the server version. A network/auth failure fails the check with the
error. In service/catalog profiles, add it as a check-only watches: entry; use
explicit checks: when a hand-written rule must share the same probe.
Response comparisons (expect). Any protocol check can assert the values
its probe returns — the server version or any field the protocol puts in its
result data (e.g. answers/rcode for dns, stratum/offset_seconds for
ntp, sys_object_id for snmp, offered_ip/lease_seconds for dhcp,
ipp_version for ipp, …). expect is a mapping of field → value (equality) or
field → {op, value} using the shared operators == != > >= < <= (numeric, or
string for ==/!=), contains (substring) and =~ (Go/RE2 regex). All
assertions must hold, in
addition to the probe succeeding:
checks:
resolver:
type: dns
host: 1.1.1.1
query: example.com
expect:
rcode: NOERROR # equality (scalar)
answers: { op: ">", value: 0 } # operator comparison
clock:
type: ntp
host: pool.ntp.org
expect:
stratum: { op: "<=", value: 3 }A referenced field the probe did not return fails the check with a clear message. The same comparison operators work for every registered protocol field.
Response latency (expect_latency). Any protocol check also accepts
expect_latency: { op, value } (milliseconds), like the http check — it fails
when the probe's response time does not satisfy the comparison. Result data
always carries latency_ms:
checks:
cache:
type: redis
password: "${env:REDIS_PASS}"
expect_latency: { op: "<", value: 50 } # alert when Redis answers slowlyGrowth bound (max_increase + within). Any protocol check can also bound
how fast a numeric field rises. max_increase maps a result field to the largest
rise allowed inside the sliding within window; the check fails while the rise
since the oldest sample in the window exceeds it. A bound of 0 fails on any
rise, which turns a since-start counter into an event: the check fails for one
within after the counter moves, then recovers. Both keys are required
together. It is evaluated after expect, so one check can hold a level and a
growth bound:
checks:
sessions:
type: redis
expect:
keys: { op: "<", value: 250000 } # level: fail above this many keys
max_increase: { keys: 20000 } # growth: fail on +20000 keys…
within: 10m # …inside this sliding window
evictions:
type: redis
severity: error
max_increase: { evicted_keys: 0, rejected_calls: 0, sync_full: 0 } # any new one
within: 15mThe first cycle only baselines, a falling value is not growth, and a field that
is missing or not a number makes the check unavailable. When no earlier sample
is left inside the window (a within no longer than the check's interval), the
previous sample is the baseline. The samples live in the check instance, so a
config reload or daemon restart re-baselines. Result data carries
<field>_increase and window.
Version-change detection (on_version_change). Set on_version_change: true
on a service check or host watch to alert when the server's version changes
between cycles — e.g. after a package upgrade. The tracked identity is the
protocol's reported version — for connection-protocol checks such as mysql,
postgres, redis, ssh, snmp, rspamd, libvirt or syncthing — or, for
protocols that only return a greeting banner (such as smtp, imap, pop,
ftp), that banner. Any registered connection protocol that reports a version or
banner participates; these names are examples, not the full set. The first cycle baselines silently;
a later change fails the check and the result data carries
version/version_old. The baseline lives in the check instance, so it persists
while the service worker or watch is alive and resets when that worker/watch is
rebuilt, for example after a config reload. It composes with on_change (the
SSH/SNMP fingerprint identity) — both can be enabled at once.
watches:
mail-version:
monitor: disabled
check:
type: smtp
host: mail.example.com
on_version_change: true # alert when the SMTP banner/version changes
expect_latency: { op: "<", value: 500 }
then:
notify: [ops-email]More protocols are added the same way — the check type, dispatch and validation are protocol-agnostic, so a new protocol only registers itself.
For an outbound mail host, declare one host watch per provider. The domain is
derived from the canary recipient and resolved through MX on every run; do not
hard-code a provider's current MX hostname. Two consecutive failures debounce a
single remote throttle. The safe example records state without sending a
notification after it is configured and enabled; replace none with an
operator-owned notifier to alert. It does not remediate because restarting the
local MTA cannot repair a remote reputation decision:
name: mail-egress-gmail
display_name: Gmail outbound acceptance
category: messaging
monitor: disabled # Replace all three identities below before enabling.
interval: 15m
check:
type: smtp_acceptance
helo: mail.sender.example
mail_from: probe@sender.example
recipient: your-owned-canary@gmail.com
starttls: required
timeout: 15s
for: { cycles: 2 }
then:
notify: [none] # Replace with an operator-owned notifier to send alerts.The clock check measures this host's wall-clock offset by querying one or more
remote NTP servers as a client. It does not require a local NTP daemon, and
the check itself never sets the system clock — checks are read-only. Correction
is a watch action: a clock watch may carry
then.makestep,
which asks the local chronyd to step the clock, or a hook running a script you own.
watches:
clock-drift:
monitor: disabled
interval: 5m
check:
type: clock
servers:
- time.cloudflare.com
- pool.ntp.org
max_offset: 2s
max_stratum: 4 # optional, default 15
max_root_dispersion: 250ms # optional
timeout: 3s
for: { cycles: 2 }
then:
notify: [ops-email]
hook:
command: [/usr/local/sbin/sermo-sync-clock.sh]
timeout: 2m
expect_exit: 0servers and max_offset are required. The check tries the servers in order and
passes when one server answers with synchronized NTP data whose absolute
offset_seconds is within max_offset, whose stratum is at most
max_stratum, and whose root_dispersion_ms is within
max_root_dispersion when that ceiling is configured. Result data carries the
selected server, port, offset_seconds, offset_abs_seconds, stratum,
leap, precision_seconds, root_delay_ms, root_dispersion_ms and
reference_id; hooks receive the same values as SERMO_* environment fields.
A failing result additionally carries clock_failure, naming which rule was
broken — offset, unsynchronized, stratum or root_dispersion — so an
action can tell a drift it can correct from a source it cannot.
A source that answers but is not synchronized — stratum 0, or a leap of
unsynchronized — always fails, whatever max_stratum allows: its offset is
near zero and would otherwise look like the best sample available.
To correct the drift rather than only report it, a chrony host can carry
then.makestep.
ntpd and systemd-timesyncd expose no step command at all, so the shipped ntpd
and systemd-timesyncd catalog services force the correction by restarting the
daemon instead, through the safe operation path. That route fires on any check
failure — an unreachable NTP server included — which is why it ships with two
servers and a 15-minute window.
source: selects where the sample comes from: ntp (the default) queries the
remote servers above, and chrony reads the local chronyd over its command
protocol. Use chrony on a host where chrony runs as a client — it serves no NTP
of its own to query, and its own tracking data is the authoritative answer for
how far this clock is from true time.
watches:
chrony-drift:
interval: 5m
check:
type: clock
source: chrony
host: 127.0.0.1 # optional, this is the default
port: 323 # optional, chronyd's command port
# socket: /run/chrony/chronyd.sock # instead of host/port
max_offset: 100ms
max_stratum: 4
max_root_dispersion: 200ms
unit: s # graph the offset over time
timeout: 3sWith source: chrony, servers is rejected (the target is the local daemon, not
a remote list) and host/port/socket address it instead; max_offset is
still required and the thresholds mean the same thing. offset_seconds is
chronyd's own correction to the system clock — what chronyc tracking prints as
"System time … of NTP time" — so it is directly comparable to the ntp source.
Result data adds the daemon's own diagnostics on top of the shared fields:
synchronized, reference_address, reference_time, reference_age_seconds,
skew_ppm, frequency_ppm, residual_frequency_ppm, rms_offset_seconds,
last_offset_seconds, update_interval_seconds and sources, sources_online,
sources_offline, sources_unresolved. As numbers in the result they are
graphable and reach hooks as SERMO_* fields. For assertions on individual
fields — a minimum source count, a skew ceiling — use the chrony check type
directly with expect:; the shipped chronyd catalog service does both.
Every type above is a single-shot check (Check.Run → Result) and is usable in
both places:
- a service's check-only
watches:entries, or explicitchecks:/preflight:referenced from rules, - a host watch document (or global
watches:entry, firing a hook) — see configuration, and - a service's own embedded
watches:block (hook/notification entries scoped to the service, or compactthen.action, including the service-scopedservice/metrictypes and the service-scopedprocess_count, which counts everything discovery attributes to the service — the init unit's control group included) — see Service watches.
Every check type is one of two styles — the style column of the type table
above is the authority, locked to the code by test. A condition check's
OK == true means there is a problem (a threshold crossed, a count exceeded),
so in rules active: {check: x} fires on it and as a watch the hook fires on
it. A health check is the opposite (OK == true is healthy), so as a watch
it fires the hook on failure; every connection-protocol check is
health-style.
A few checks report an increment rather than a state: oom reads the kernel's
cumulative oom_kill counter and reports the per-cycle delta, so a kill makes the
condition true for exactly one cycle and false again on the next.
A for:/within: window asks the condition to hold for several consecutive cycles.
On an edge condition that can never happen, so the window silences the sensor
entirely — the watch stays green right through the event. Configure an edge sensor
with no entry window at all, so it fires on the first increment:
name: watch-oom
interval: 30s
check: { type: oom }
clear: { duration: 1h } # how long the alert stays up; see belowWithout an entry window the alert would also end on the next cycle, one 30s flash
nobody sees. That is what clear: is for — it holds the firing episode open while the
condition is quiet, so an edge sensor's whole visible alert life is its clear window.
The lifecycle:
| Cycle | What happens |
|---|---|
| the kill | condition true → fires immediately → firing event + one notification |
| next cycles | delta back to 0; the clear window holds the alert up; no repeat actions |
after clear: |
recovered event → sensor back to ok |
Notifications are sent once per episode, on the rising edge; add
then.notify_interval if you want reminders while the alert is up. A later increment
after recovery opens a fresh episode and notifies again.
Sermo warns at startup when a watch gates an increment-only check behind a window it
cannot satisfy, naming the watch and the ceiling. It is a warning and not an error so
an upgrade is never blocked by a config that predates the check. Rate-shaped deltas
(swap io, net errors) are not edge sensors — they stay true while the pressure
lasts, so a window is legitimate hysteresis for them; a count delta likewise holds
for its own within span and is only flagged when the rule window outlasts it.
The multi-metric watches (net, icmp, swap) keep their metrics: map shape
(one hook per metric) watch-only, but their single-metric form — an explicit
metric: field producing one result, e.g. {type: net, interface: ppp0, metric: state, expect: up} or {type: icmp, host: 1.1.1.1, metric: state, expect: up} —
works as a service check-only watch or explicit checks: entry (used by the
pppd catalog daemon to watch its uplink). The multi-target watches (file, process, one
event/hook per changed path or matching pid) stay watch-only.
service/metric/process checks need per-service context (backend status, a
metric sampler, process discovery) and so are not available as standalone
watches.
The route check verifies the kernel has an up default route — read
natively through netlink from the IPv4 (the default) or IPv6 (family: ipv6)
routing tables. A route or multipath hop the kernel flags linkdown (the
interface is up but has no carrier) or dead does not count. With
interface, a default route must egress through that interface. It is a
health check (OK means the route is there); as a watch it fires when the route
disappears or its link goes down.
It closes the uplink gap the link and ping layers leave: after a failed PPP
renegotiation the interface can stay up with the default route gone, and a
ping bound to the interface cannot tell "no route" from "provider down". The
pppd catalog service layers all three (net state, route, icmp).
checks:
route:
type: route # IPv4 by default; family: ipv6 for the v6 table
interface: ppp0 # optional: the default route must leave through ppp0The result reports the matched egress interface and gateway (when the route
has one — point-to-point links have none) in its data, and value carries the
number of matching default routes.
The firewall_rules check verifies that nftables or iptables rules are loaded.
It is health-style: a service or watch fails when the rule count is below
min_rules (default 1). backend: auto tries nftables first (read via
netlink, no nft binary required) and falls back to iptables/ip6tables.
checks:
service: { type: service, expect: active }
firewall:
type: firewall_rules
backend: auto # auto | nftables | iptables
min_rules: 1
requires: [service] # useful for oneshot firewall loadersAs a watch, it fires the hook when the firewall rules disappear. Hook extras:
SERMO_BACKEND, SERMO_RULES, SERMO_MIN_RULES.
The failed_units check counts the init units the host reports as failed and
names them. It exists because service monitoring only covers configured
services: a site-local unit with no catalog profile — a nightly backup job, say —
is otherwise invisible, and so is a failed .mount or .timer. The listing is
therefore not restricted to .service.
watches:
watch-failed-units:
category: system
interval: 1m
check:
type: failed_units
backend: systemd # optional: auto (default) | systemd | openrc
count: { op: ">", value: 0 } # optional; this is the defaultbackend: auto detects the init system on every cycle, so a generated
configuration names the host's real backend instead. On systemd the units come
from systemctl list-units --state=failed; on OpenRC from the crashed services
in rc-status --all.
The check is condition-style and has no remediation action: restarting an
arbitrary unit Sermo knows nothing about is not a safe action, so a failed unit
is reported and left alone. Data keys are backend, count and units (the
joined unit names), and the dashboard shows all three.
fds, pids and conntrack measure a count against a kernel ceiling. Two
kernels give them nothing to measure against: one that reports no ceiling, and
one that reports an unreachable one — fs.file-max reads 9223372036854775807
on a host that lifts the cap, and 879116 descriptors of that is 0.0%.
Neither is a denominator, so neither produces a percentage. The check publishes
its count and says why the rest is missing — fds 879116 allocated (no kernel limit), or (limit unknown) — and omits used_pct, free and the ceiling
itself rather than reporting them as zero. The count is what its scalar carries,
the panel shows it as a value instead of a gauge, and no flat 0% series is
recorded.
A ceiling also has to be the one that binds. pids counts threads, and both
kernel.pid_max and kernel.threads-max cap them; it measures against the
smaller. Dividing by pid_max alone reports whichever is looser — on a host with
pid_max 4194304 and threads-max 1027204 that understates utilisation fourfold,
so a table a quarter full reads as one-sixteenth full.
The consequence is worth stating plainly: a used_pct threshold cannot fire on
such a host. It never could — it just used to look green rather than say so. To
watch descriptors there, write the threshold against the count:
watches:
watch-fds:
check:
type: fds
allocated: { op: ">", value: 2000000 } # a count, not a share of nothingThe inotify check reports the two per-user kernel limits,
fs.inotify.max_user_instances and fs.inotify.max_user_watches, for the user
closest to each of them.
It exists because no other check can see this exhaustion. fds compares
system-wide allocated descriptors against fs.file-max, which is effectively
unlimited on a modern kernel, so a host whose uid 0 held all 1024 inotify
instances — every new login failing to start a user manager, every new session bus
failing to initialise inotify, systemctl is-system-running reporting degraded
— showed nothing at all in fds. It no longer reports that as 0.0%; see
An unreachable ceiling is not a ceiling.
watches:
watch-inotify:
category: system
interval: 1m
check:
type: inotify
used_pct: { op: ">=", value: "80%" } # the worse of both limits
# instances_used_pct / instances_free / instances
# watches_used_pct / watches_free / watches
for: { cycles: 3 }used_pct is the worse of the two utilisations, and it is the field the
generated watch uses. Level predicates are ANDed, so a check carrying both
instances_used_pct and watches_used_pct would have stayed silent on the host
above: 100% of the instance limit and 0.7% of the watch limit. The prefixed
fields are there for an operator who knows which limit they care about — a build
host that legitimately holds many watches can threshold instances_used_pct
alone.
Both limits are charged per user, so the check reports the worst single user
per limit, and the two can be different users: root leaking instances while a
desktop user holds every watch. An aggregate would report 120% while nobody is
anywhere near being denied an inotify_init.
Counting instances costs one readlink per open descriptor. Counting watches is
gated on a watches/used_pct predicate being present, and then costs one
streamed fdinfo read per descriptor already known to be an inotify one — so it
is proportional to inotify descriptors, not to all descriptors. Data keys are
instances, instances_max, instances_uid, watches, watches_max,
watches_uid, dimension (which limit is binding), users, holders (the
process names holding the instances, which is usually the whole diagnosis) and
unreadable.
The check needs root: /proc/<pid>/fd is owner-only, so an unprivileged run
counts only its own processes. Rather than report a reassuring ok from a partial
walk, it reports how many fd tables were unreadable and marks the sample a lower
bound.
The hdparm check times a disk's read throughput and alerts when it crosses a
threshold — useful to catch a gradually degrading drive. It runs hdparm on
device and exposes two MB/s values: read (buffered disk reads, hdparm -t
— the real device speed) and cached (cached reads, hdparm -T — memory/cache
throughput). Predicates are {op, value} in MB/s; at least one of read/
cached is required, and only the timings a predicate needs are run (a
cached-only check skips the slow buffered pass).
watches:
disk-speed:
interval: 24h # hdparm -t reads ~3s and adds I/O — run rarely
check:
type: hdparm
device: /dev/sda
timeout: 30s # give the benchmark room
read: { op: "<", value: 100 } # alert when buffered reads drop below 100 MB/s
cached: { op: "<", value: 3000 } # optional: cache/memory throughput
then:
notify: [ops-email]hdparm is condition-style: a predicate expresses the alerting condition
(e.g. read < 100), so the watch hook/notify fires when it holds.
hdparm needs root (raw device access); without it the
check fails with hdparm's error. A run that exits non-zero, or that lacks a
timing a predicate asks for (a disk that fell off its bus still answers -T
from cache before -t fails), is an unavailable observation, never a sample
judged on the half that came back. Because -t reads from the platter for a few
seconds and adds real I/O load, schedule it on a long interval (e.g. 24h)
with a generous timeout. The measured read/cached are placed in the result
data (and the SERMO_READ/SERMO_CACHED hook variables), and are recorded as a
time series and graphed in the service detail (web UI) so you can spot gradual
degradation of a drive over time. (This per-check named-metric graphing is
generic: any check that publishes numeric Result.Data fields can opt in.)
A disk that has spun down cannot be benchmarked. hdparm -t wakes it, the
spin-up eats the whole timing window, and the drive reports a single block in
several seconds — a fraction of a MB/s, which hdparm prints as kB/sec instead
of MB/sec. Sermo reads both spellings and always records MB/s, so that sample
is a real, very low read and will cross a read < threshold. Point a
throughput watch at disks that stay awake, or grade it severity: warning (see
Severity) so the drive's own power management reads amber
rather than red.
smart, hdparm and diskio publish the transport their device sits on —
sata, usb, nvme, scsi, virtio, mmc or virtual — read from the sysfs
path the kernel links the device to. It explains behaviour the numbers alone do
not: a USB disk parks itself, so it answers a throughput benchmark with its own
spin-up rather than with its media speed.
smart, hdparm and diskio address a block device by name, and a disk that
dies rarely disappears: it keeps its /dev node and its /proc/diskstats row
and simply stops answering. Left alone, each check reads that as good news —
smartctl returns a report with no verdict, hdparm an unreadable timing,
diskio a window with no I/O — so a dead drive would show as healthy.
All three therefore classify the device before they classify the reading. When
smartctl reports it could not open or identify the drive, or sysfs no longer
sizes the device (/sys/class/block/<device>/size gone or 0), the check
publishes missing: an unavailable observation, never a passing one. The
dashboard shows missing as the watch state and health, and the row reads as a
failure. diskio consults sysfs only for a window that moved nothing at all, so
a disk carrying traffic never pays for the lookup. The row is not left otherwise
blank: see What a device that stopped answering still
reports.
A missing device is unavailable, not firing, so — like every unavailable
observation — it never triggers automatic actions. It is reported through the
watch's error event and the dashboard.
An unavailable sample is the moment an operator most needs to know which disk
this is, so smart, hdparm and diskio do not fall back to the word
missing alone.
Identity comes from sysfs, which keeps publishing what the kernel learned
from the drive's last successful identification long after the drive stops
answering: model, serial_number, firmware, wwn (the World Wide Name, as
/dev/disk/by-id spells it) and rotation_rate, the medium the drive uses.
capacity_bytes is the exception — sysfs sizes a dead device at 0 — so it is
reported only while the device is still sized. A
smart sample that did reach the drive uses smartctl's own richer answer
instead, which is where a full serial number comes from: sysfs truncates a SATA
model to 16 characters and publishes no serial for it at all.
Last known readings come from the check's own memory of the newest sample
the device answered: last_health and one last_<field> per reading, dated by
last_seen_seconds. They are deliberately published under their own keys and
labelled (last) in the dashboard, never under the live key, because the
recorder graphs every numeric a result carries — republishing a dead drive's
final temperature as temperature would draw a flat line at it forever.
That memory belongs to the running check, so it is empty after a daemon restart: Sermo reports only samples it actually took. Identity survives a restart because the kernel, not Sermo, is the one remembering it.
A network interface is the exception that proves the rule. Unlike a disk,
which keeps its /dev node and its sysfs identity after it goes quiet, an
interface that is removed or renamed takes its whole /sys/class/net directory
with it, so there is nothing left to read. A net check therefore remembers the
identity it last observed and reports that: mac, driver, bus (the device's
address in its bus tree, which is what distinguishes one port of a multi-port
card), mtu, duplex and — for a virtual interface, which has no driver to
name — kind, the bridge/vlan/bond/veth the kernel calls it. A live
sample carries the same rows plus carrier_changes, the kernel's count of every
link transition since the interface appeared: a link that is up now but has
flapped two hundred times is not the same situation as one that has been up since
boot, and a check sampling once a cycle cannot see the flaps between its own
samples.
Physical-health checks are condition-style: predicates are alerting conditions. Numeric values are recorded over time and graphed in the service detail so gradual degradation is visible.
-
sensors— lm-sensors-style hwmon inputs (no external tool; reads/sys/class/hwmon). Aggregates:temp(the hottest matching temperature, °C),fan(the slowest matching fan, RPM — catches a stalled fan) andvoltage(the lowest matching rail, V). At least one predicate is required; optionalchipandlabelsubstrings narrow which inputs count. An input that measures nothing is left out: a channel its driver disabled (_enable0) or flags faulty (_fault1), and a temperature outside -55..125 °C — the saturated 127 or -128 °C an unused probe input (a board's AUXTIN) reports.checks: cpu-temp: type: sensors chip: coretemp # optional: only this chip temp: { op: ">", value: 85 } # alert when the hottest core exceeds 85 °C fan: { op: "<", value: 400 } # optional: alert on a stalled fan
-
smart— a drive's SMART health viasmartctl -i -H -A -c -l selftest -j(needs smartmontools and root). With no predicate it alerts when the overall SMART verdict is FAILED; predicates addtemperature(°C),reallocated(sector count, an HDD failure sign),pending_sectors(sectors the drive could not read and has not managed to reallocate — the count that rises beforereallocateddoes),crc_errors(corrupted transfers on the link itself, which blames the cable or the backplane rather than the media),media_errors(the NVMe counterpart ofreallocated: NVMe drives publish no attribute table),wear(SSD/NVMe percentage used) andpower_on_hours. Complementshdparm(throughput) with failure prediction. The drive's FAILED verdict and every configured predicate are independent alert conditions: any one is enough. They are not graded alike: a predicate holding under a PASSED (or unknown) verdict is awarning— amber, out of aggregate health and the SLA, notified as[sermo][warning]— while the FAILED verdict and an unreadable drive areerror, unless the check declaresseverity:(see Severity). Thehealthreading keeps sayingPASSEDbeside the amber row, which is exactly the situation it describes. A missing transport-specific field never suppresses another field, so an ATA drive can alert onpending_sectorswhile omitting NVMe-onlymedia_errors.Every sample also carries what the drive is —
model,serial_number,firmware,wwn,capacity_bytesandrotation_rate— pluspower_cyclesandself_test, the drive's own verdict on the last self-test it ran and the lifetime hour it ran at. Those name the disk an operator has to pull out of a bay, which a device node does not.rotation_rateis the drive's own figure (7200 rpm, orSSDfor flash); an NVMe drive reports none, so the kernel's coarser answer (rotationalorSSD) stands in. It is what makes the rest of the readings mean something:wearis the number that matters on flash,reallocatedon platters. A drivesmartctlcannot open or identify is reported missing and unavailable — see Missing devices.checks: ssd-health: type: smart device: /dev/nvme0 interval: 1h reallocated: { op: ">", value: 0 } # any reallocated sector pending_sectors: { op: ">", value: 0 } # unreadable sectors awaiting reallocation crc_errors: { op: ">", value: 0 } # link errors: suspect the cable wear: { op: ">", value: 90 } # SSD/NVMe nearly worn out
-
storcli/ssacli— read-only hardware-RAID health. These are host watches, not service-process checks: the utilities do not run as daemons and the controller hardware belongs to the host.binarymust be an absolute executable path. StorCLI reads controller/enclosure, physical-drive and virtual-drive JSON reports; SSA CLI readsctrl all show config detail. Sermo alerts on unhealthy controllers, arrays/volumes, drives, enclosures, cache or battery/capacitor protection; preserved offline cache, safe mode, shutdown requirements; unrecoverable media, memory, other or predictive errors; SMART alerts/wearout; inconsistent or inaccessible volumes; and incomplete parity/rebuild work. The findings are graded: a state finding (any of those components not OK, memory errors, a drive's SMART alert or wearout, unfinished parity or rebuild work) is anerror, while a counter finding on members whose state is still OK — media, other or predictive error counts, a volume reporting unrecoverable media errors, thetemperaturepredicate — is awarningunless the check declaresseverity:. Thehealthreading saysok,warningorerroraccordingly,raid_issueslists every finding andraid_advisoriesthe counter subset. Each controller, cache, virtual volume and physical drive is exposed as its own reading. Controller readings include model, firmware and on-board RAM/cache; volume readings include RAID level, capacity, Linux device, cache policy and active operation; drive readings include bay, state, capacity, media/interface, model, serial, firmware, temperature, error counters and the controller's physical-drive SMART verdict. Rebuild/reconstruct progress is attached to the affected volume or drive and also exposed asraid_progress_pct.temperatureoptionally alerts on the hottest controller, cache-protection module or drive; a report with no temperature at all publishes notemperaturereading (the message saysmax_temperature=unknown) and the predicate does not hold. Counts, reasons and the maximum temperature are exposed as readings, and temperature/error/progress history is graphed. Do not add asmartwatch for an OS device that is one of these virtual volumes: smartctl sees the logical volume, while this controller watch is the source of truth for the member drives.checks: controller: type: storcli # use ssacli for HPE Smart Array binary: /usr/bin/storcli64 timeout: 2m temperature: { op: ">", value: 70 }
Sermo never invokes configuration, rebuild, patrol-read or write commands. Fleet generation creates a watch only after the installed utility answers a read-only controller query, so a vendor CLI package with no matching controller does not create a permanently unavailable watch.
-
raid— Linux md software-RAID from/proc/mdstatand read-only/sys/block/md*/mddata (native). With no predicate it alerts when any array is degraded; aninactivearray (listed but not started, so its data is unavailable) counts as degraded and readsinactive. Predicates adddegraded,recoveringandarrayscounts.array: md0scopes the check to one array. Withsysfs_changes: true, Sermo tracksmismatch_cntand each member'sstate,errorsandbad_blocksbetween cycles. A host with no md arrays never alerts.checks: raid: { type: raid } # alert if any md array is degraded
A RAID host watch can filter its
then.notifytargets to lifecycle transitions withthen.notify_on:on_degraded,on_recovering,on_good(repair complete) andon_array_change. These notifications receiveSERMO_RAID_EVENT,SERMO_RAID_ARRAY, operation/progress and, for sysfs changes, member, field, old and new values; notifier templates can use those fields to render a different message.A host watch can additionally opt into manual reconstruction pause/resume with
raid_control: { pause_resume: true }; see Manual RAID reconstruction control. -
lvm— Linux LVM health and capacity from read-onlylvsJSON. It is a health check:okmeans the selected VG/LV is usable and its configured limits are respected. A breachedfree_pctthreshold alone reportswarningby default, even at 0 %: VG allocation headroom is not filesystem free space.errorcovers an absent, partial or suspended LV, a reported LV health fault, or a configured thin-pool capacity threshold. Explicitseverity:overrides the default grade without changing the raw check verdict. Select a target withvolume_groupand optionallogical_volume;free_pct,thin_data_pctandthin_metadata_pctare ordinary numeric predicates. Withoutlogical_volumethe watch covers every LV of the group, hidden ones included: any faulty LV or any thin pool over its threshold fails it,lvm_reasonsnames each faulty LV (data:partial), and the thin-pool readings report the fullest pool. Result readings includehealth,volume_group,logical_volume,lvm_reasons,vg_free_bytes,vg_size_bytes,vg_used_bytesand the configured percentage fields.watches: lvm-vg0-root: check: type: lvm volume_group: vg0 logical_volume: root thin_data_pct: { op: ">=", value: "80%" } then: notify: [ops] notify_on: [on_change]
on_changeis LVM-only and sends one notification when the effective health changes betweenok,warninganderror, including an advisory escalating to a volume fault; it does not notify repeatedly while the same state persists. Templates receive VG/LV, current and previous states, current reasons and recovered reasons. Panic mode suppresses that delivery and records the state transition as apanic-suppressedevent. -
edac— ECC memory errors from the kernel EDAC subsystem (native,/sys/devices/system/edac).ceis the cumulative correctable count anduethe uncorrectable count; with no predicate it alerts onue > 0. The check fails when the platform exposes no EDAC controllers (so you notice ECC isn't reported).checks: ecc: type: edac ce: { op: ">", value: 100 } # also alert on many correctable errors
Numeric resource and counter predicates ({op, value}) require finite values;
NaN and infinities are rejected during configuration validation as well as
check construction. Percentage predicates remain bounded to 0–100, and byte
predicates require an explicit size suffix.
A count check tallies the entries in a directory and either compares the total
to a threshold, or alerts when the total grows by a delta within a time
window. It is condition-style (OK == true means the comparison holds), so
in rules active: {check: …} fires when the comparison holds and
failed: {check: …} fires when it does not.
checks:
spool-backlog:
type: count
path: /var/spool/myapp # required: directory to scan
of: file # any (default) | file | dir | symlink
recursive: false # optional, default false
include_hidden: false # optional, default false for recursive scans
op: ">" # >=, >, <=, <, ==, !=
value: 1000 # numeric thresholdchecks:
spool-growth:
type: count
path: /var/spool/myapp
of: file
delta: { op: ">", value: 200 } # alert if the count grows by >200…
within: 2m # …within this sliding windowofselects which entries are counted. Entries are classified by their own type without following symlinks, so a symlink counts assymlink(never as the file or directory it points to);anycounts every entry.recursive: truedescends the whole subtree (the directory itself is never counted); unreadable subdirectories are skipped. Hidden descendants (names starting with.) and their subtrees are skipped by default; setinclude_hidden: trueto count them. Default counts only the immediate entries.- A missing or unreadable
pathmakes the check fail. The observed total is exposed in the check's result data ascount. - The threshold may also be written as a nested predicate —
count: { op: ">", value: 1000 }— matching the{op, value}form the other checks use. Use one form or the other, not both. delta+withinis stateful. Each cycle samples the count, keeps samples in the lastwithin, and compares the current count against the oldest sample still in the window. When no earlier sample is left in the window (awithinno longer than the check's interval), the previous sample is the baseline, so growth is measured cycle to cycle instead of reading+0forever; the message then shows the real, slightly longer span. The first cycle only baselines (no alert), and only increases can trip the check; steady or shrinking directories pass. Result data carriescount,baseline_count,growth_count,windowandvalue(the growth). Use eithercount/op/valueordelta/within, not both.
A log check follows a log file like tail -f and counts the lines appended
within a sliding window that match a regular expression. It is
condition-style (OK == true means the comparison holds), so a watch with
count: { op: ">", value: 3 } fires its then: when more than three matching
lines arrived within within, and releases once the window has slid past them.
It exists for the failure a service reports only in its own log: a consumer
whose telemetry exporter times out on every message
(cURL error 28: Connection timed out) keeps its init unit active and its
process alive while it crawls.
checks:
otlp-export-timeouts:
type: log
path: /var/www/app/current/var/log/symfony-messenger_*_err.log # absolute; a glob sums its files
regex: 'cURL error 28: (Connection|Operation) timed out' # Go/RE2, matched per line
count: { op: ">", value: 3 } # matching lines within the window
within: 5m
optional: truepathmust be absolute and point to regular files. FIFOs, sockets, directories and devices are rejected, including through symlinks; opening a FIFO does not wait for a writer. A glob (*,?,[) matches files whose matches are summed; the result data reportsfiles. Keep the glob tight (*_err.log, never*_err.log*): a rotated copy that matches the glob is a new file and would be read from its start.- The first cycle only baselines at the end of every file: lines already there are history, not news. From then on each cycle reads what each file gained. A trailing partial line waits for its newline.
withincounts observation time, not timestamps embedded in log text. Each batch stays in the window from the cycle that read its complete lines. New matches still count after a pause longer than the window, or when the check interval exceeds it; matches from older observations expire normally.- Rotation is handled by identity and size: a new inode under the same name
(logrotate's rename + create) or a file that shrank (
copytruncate,truncate) is read again from its start, and a file that vanished is forgotten. - Read budget: one cycle reads at most 8 MiB across the matched files,
including bytes of partial lines that must be re-read next cycle. Past
that the remainder is skipped, the result carries
truncated: trueand the message says the count is a lower bound — a log that explodes must not stall the worker. - State is in memory. Like
count'sdeltaandsize, the offsets and the window live in the built check:sermoctl daemon reloador a restart re-baselines, and lines written meanwhile are not counted. - Result data carries
count(matches within the window, also the graphedvalueinlines),regex,window,files,bytes_readandtruncated. A missing or unreadable file makes the check unavailable; withoptional: truethat is a warning rather than a failure.
Service metrics measure the discovered process set; system metrics measure the
machine. value is a number with an optional trailing %.
scope: service memory, swap, cpu, cpu_thread, process_count, io, io_read, io_write, fds, threads
scope: system total_memory, total_swap, total_cpu, load1, load5, load15
A scope: system metric may only drive alert rules, never remediation. It
describes the whole machine, not one service, so a restart/start/stop rule
that reads a system metric — directly, or through a failed/active reference
to a type: metric, scope: system check — is dropped at config load with a
warning. This is a safety invariant: a system-wide signal must never act on an
individual service (see docs/safety.md).
Service metrics sum across the whole discovered process tree — the matched
processes and their child/descendant processes — so a service's cpu,
memory, io, fds, etc. account for its workers and helpers, not just the
main process. io/io_read/io_write are byte/second rates over actual
block-layer I/O (io is read+write); fds is the open file-descriptor count
(summed) and threads the thread count. The fds percentage is the one
per-process figure: the process of the tree that is closest to its own soft
RLIMIT_NOFILE (/proc/<pid>/limits, "Max open files"), because the limit is
per process and the process about to hit EMFILE is the one that stops
accepting connections whatever the rest of the tree holds.
The rates (cpu, io, io_read, io_write) add up each process's own
change since the previous sample, counting only processes seen in both samples.
A child that exits between cycles (php-fpm pm.max_requests, Apache
MaxConnectionsPerChild, Postfix smtpd) therefore does not subtract its
lifetime totals from the rest of the tree, and a process that missed one sample
does not come back as a spike. The work a process did in the interval in which
it started or exited is not counted.
memory is the summed RSS (resident memory) of the process tree, as bytes
and as a percentage of total RAM. swap is the summed swapped-out memory
(VmSwap) of the tree, as bytes and — when a swap device exists — as a
percentage of total swap; it is reported only on hosts where swap accounting is
readable.
The cpu percentage is the service's summed CPU time (parent + children) over
the elapsed wall-clock, normalized by the server's total logical CPUs (the
hardware threads, counted from /proc/stat so the figure reflects the whole
machine even if Sermo is pinned to a CPU subset). So 100% means the service's
processes are saturating every CPU thread of the server, and a single fully-busy
core on an 8-thread host reads ~12.5%. total_cpu uses the same whole-machine
basis: the non-idle share of the aggregate /proc/stat line, where virtual
machines' guest time is counted once, inside user time, as top does.
cpu_thread complements cpu for the single-thread case: it is the busiest
single thread in the tree (of the parent or of any child) measured against one
CPU thread, so 100% means one thread is saturating a full core and the metric never
reads above 100% — no single core can give more. Because the whole-machine cpu
dilutes one hot thread across all cores (a core-bound process on an 8-thread host
shows only ~12.5% there), cpu_thread is what you alert on to catch a thread
pegging its core: metric scope: service, metric: cpu_thread, op: ">", value: "90%". cpu_thread is a rate, so it is not ready on the first cycle.
Reading a process's individual threads costs one file per thread per cycle, so it is done only for processes at or above 5% of one core. Below that the process's own rate is published as an upper bound: no thread can exceed the whole process, so a saturated core can never hide behind a bound — it can only over-report, and only far below any threshold worth alerting on. Keep rule thresholds above that floor; a threshold under it would compare against a bound rather than a measurement. A process that has just become busy is bounded for one cycle and measured from the next, because a per-thread rate needs two per-thread samples. The Web UI's process table shows the per-process figure, with its tooltip saying whether that figure was measured per thread or bounded by the process rate. The cell itself carries no marker: below the floor a bound and a measurement are indistinguishable for any decision, and on an idle host every non-zero row is a bound — a marker on every row distinguishes nothing.
cpu/cpu_thread/total_cpu and the io* metrics are rates: they are not
ready on the first cycle and a condition over a not-ready value is false. A %
threshold needs a metric with a percentage form (memory, swap, cpu,
cpu_thread, fds, total_memory, total_swap, total_cpu;
swap/memory/fds/total_memory/total_swap also have an absolute form); a bare number needs an absolute form (everything else, including
io*/threads, which are absolute only). Reading another process's I/O, fd
count or limits requires privilege, so those sum only the processes the daemon
can read, and the fds percentage is absent when no limit could be read or the
limit is unlimited.
Because io*/threads have no percentage form, a packaged threshold for them
cannot be normalized to the host the way memory and cpu_thread are, and a
service-scoped metric sums every process discovery attributes to the
service, control-group members included. A catalog default is therefore a
deliberately generous ceiling rather than a tuned number, and the host that needs
a different one says so in services.local/ (see
per-host overrides). fds is
the exception: its percentage is measured per process against that process's
own limit, so the restart-if-fds-high rule Sermo injects into every service
uses 80% (fds_limit) as a default alert threshold. Restart requires
restart_on_fds_high: true. Tune the threshold against representative load:
closeness to the limit is not proof of a leak. An absolute ceiling such as
50000 can never fire for a daemon whose
limit is 32768, which is how a collector leaking one socket per accepted
connection reached 32761/32768 unnoticed. Where the control group holds
workload the daemon does not own — a hypervisor's per-domain helpers, a
container runtime's containers — a summed absolute count describes that
workload rather than the daemon, so the catalog ships no absolute fd watch for
such services and host-wide exhaustion is alerted from a host watch instead.
Rule evaluation is deterministic and order-independent: guards always run before remediation, at most one remediation action runs per service per cycle, and when several remediation rules fire at once they are considered in sorted name order — the first non-blocked action wins. Every declared check and inline condition probe runs at most once per cycle; rules read the cached results.
rules:
RULE_NAME:
type: remediation | guard | alert
if: { ... } # condition tree
for: { cycles: 3 } # consecutive cycles (optional)
# for: { duration: 6m } # or consecutive wall-clock time
within: { cycles: 15, min_matches: 5 } # sliding window (optional)
# within: { duration: 30m, min_matches: 3 } # or a time window
clear: { duration: 6m } # recovery hysteresis, alert rules only (optional)
notify: [ops-email] # who gets this rule's alert messages (optional)
emission: { events: on_change, notify: on_change } # or every_cycle
then: { action: alert, message: "http is down" }A rule's notify selects which notifiers receive its alert messages,
overriding the global default (Notifications):
an explicit list wins, notify: none suppresses, and omitting it inherits the
global notify default. It applies to the rule's alert messages; remediation
operations are reported as events, not notifications. By default those automatic
alert events and notifications are emitted only when the rule enters a firing
episode, then recovered is emitted when it clears; remediation rules emit
that recovered event too. Use rule-level
emission.events or emission.notify (on_change | every_cycle) to override
the global emission policy for that rule. Operation result events remain audit
events and are recorded whenever the operation is attempted. While panic mode
suppresses an operation, its accompanying alert follows the configured emission
policy: on_change records episode transitions and severity escalation without
repeating the same alert each cycle. No action cooldown is recorded until an
operation is attempted.
The separate fleet-wide event_notify route
can send these alert and recovery events even when the service is in dry_run.
It also covers service check failures that have no notification rule.
For a recovered rule with exactly one direct check or metric leaf, the event also
records the current formatted value and its configured operator and threshold.
Byte sizes use IEC binary units — B, KiB, MiB, GiB or TiB — matching
how size suffixes are parsed in configuration, including the configured
threshold; byte rates use SI decimal units — B/s, KB/s, MB/s, GB/s or
TB/s. This makes threshold flapping visible without having to reconstruct the
sample from the metrics history.
Every operator-facing number follows one canonical convention on every surface
(events, notifications, sermoctl and the web UI): comma as the thousands
separator and dot as the decimal mark (12,345.68), with byte sizes always
humanized to IEC units (2.44 MiB) and byte rates to decimal units
(1.02 KB/s) by the same formatter. Durations render as space-separated whole
components, greatest-first with zeros omitted (2h 3m 20s,
3y 2mo 10d 12h 20s), promoting the head unit only well past its boundary:
bare seconds up to 360s, hours up to 72h (70h 15m), days up to 120d, months
(30 days) up to 24, then years (12 such months). Rule-window progress
(3m/6m) is the one exception: it echoes the operator's configured window in
its own spelling. Event
timestamps are always rendered in UTC, so the same instant reads identically
before and after a daemon restart.
Actions and types are coupled: the operation actions (restart, start,
stop, reload, resume) belong to type: remediation rules — required there (a
notify-only rule is type: alert) and rejected elsewhere. alert (with a
message) may accompany any rule's actions; block is guard-only. A then
may carry one action or an actions list (e.g. alert + restart together).
Those operations use the same safety engine as manual CLI/Web actions, including
the active-service exact process identity gate before restart.
repair is deliberately not a rule action. It is available only to an operator
through sermoctl repair and the dashboard for a failed or inactive service
whose residual runtime or init state needs the extra, guarded recovery steps;
normal failure remediation uses restart.
Conditions form a logical tree with and/or/not and leaves:
if:
or:
- failed: { check: http } # a named check failed
- active: { check: backup-flag } # a named check passed
- file: { path: /run/x, exists: true }
- command: { user: postgres, command: ["/usr/local/bin/can-restart", "${service}"], timeout: 5s, expect_exit: 0 }
- service: { state: active }
- process: { exe: /usr/bin/mysqld, user: mysql, state: running }
- metric: { scope: service, name: cpu, op: ">", value: 30% }
- changed: { path: /lib64/libc.so.6 } # the file changed since the last cycle
- changed: { app: containerd, level: patch } # app version changedcommand is a direct condition leaf whose truth is the same as a command check:
exit status/output expectations pass. It must use array argv form and declare a
timeout; user is available with the same meaning as on a command check. It is
run without a shell, cached for the cycle like other inline probes, and must be
side-effect-free. failed/active may also take an inline
probe (tcp, command, ...) instead of a check: reference when you need the
named success/failure polarity.
changed is true when the file at path differs (size/mtime) from the baseline
tracked across artifact samples, or when app names a linked app whose version
changed at the selected level (major, minor or patch; default patch).
The first cycle adopts the current value (a daemon start never fires), and a
successful restart/start re-baselines it. A failed app version command is an
invalid sample: it does not fire and does not update the baseline. An app that
is not installed (including a failed version_match) is an unavailable sample:
it is false without a rule-evaluation error and does not set a baseline. When it
is later installed, its first version establishes that baseline. This does not
relax the linked app's normal operation preflight. The path form is the
primitive behind restart_on_change.libraries (see Services →
Library services); the app form is the primitive behind
restart_on_change.apps for service-owned binaries such as containerd.
For service paths, linked apps and catalog libraries, the samples run at
engine.artifact_interval (or their local interval), not on every service cycle.
Without for/within, a rule fires the cycle its condition is true.
for is consecutive: for: {cycles: N} requires N consecutive true cycles,
while for: {duration: 6m} requires the condition to stay true for at least
that wall-clock duration. within is a rolling window:
within: {cycles: N, min_matches: M} requires M true cycles out of the last N,
while within: {duration: 30m, min_matches: M} requires M true observed cycles
inside the last 30 minutes. min_matches is optional and defaults to 1
(true at least once within the window). A rule cannot use both for and
within; a single window must choose either cycles or duration, not both.
clear is the recovery-side window (anti-flapping hysteresis): once the rule is
in a firing episode, clear: {cycles: N} or clear: {duration: 4m} keeps the
episode open until the condition has stayed false for the whole window, so a
metric oscillating around its threshold produces one episode (one alert, one
recovered) instead of one per crossing. A cycle where the condition is true
again resets the clear progress and continues the same episode without a second
alert. A clear window only extends an episode — it never cuts the entry window
short. clear is only valid on type: alert rules and on watches: a
remediation or guard rule held firing on a false condition would keep acting on
it, so validation rejects it there.
Every alert rule and watch has a clear window: when the target declares no
clear: of its own, the clear_window: {cycles | duration} fallback applies —
configurable in global defaults and per service, like rule_window — and
without that the built-in default is 5 minutes. clear: {cycles: 1} (or a
clear_window of one cycle) opts a target back into immediate clearing: one
false cycle ends the episode.
Service rule-window progress is persisted in paths.state. If sermod
restarts while a for window is at 2/3 consecutive matches, the next observed
matching cycle continues from 2/3 instead of starting from zero. Duration-based
windows persist their timestamps too, so a restart does not restart a pending
for: {duration: ...} window. The firing episode and clear progress persist the
same way, so a restart neither re-alerts an open episode nor drops a pending
clear window.
Guard rules block unsafe actions and use action: block with a message:
block-during-backup:
type: guard
blocks: [restart, stop]
if: { file: { path: /run/backup/in-progress, exists: true } }
then: { action: block, message: "Backup is running" }Guard rules evaluate their condition at the moment an action is requested, so
they do not support for: or within: windows; declaring either on a guard is
a validation error.
blocks: accepts restart, start, stop, reload, resume and the
manual-only operations repair, reap (sermoctl reap --apply),
close_session (closing an SSH or terminal session),
close_terminal_source (closing an empty tmux server) and kill_query
(cancelling a database statement listed by a db_queries
watch, manually or through its opt-in then.kill_query). Any other value is a
validation error, so a typo cannot silently disable a guard. A guard applies to
the action it names and to every action that performs that step:
blocks: entry |
also denies |
|---|---|
start |
restart, repair |
stop |
restart, reap |
Session closes signal one session process and do not start or stop the service,
so only close_session / close_terminal_source deny them. Likewise only
kill_query denies a statement kill:
rules:
no-kills-during-batch:
type: guard
blocks: [kill_query]
if: { file: { path: /run/batch/nightly.lock, exists: true } }
then: { action: block, message: "nightly batch running; statements are not cancelled" }Connection thresholds are workload policy, so catalog services do not enable them by default. Add an explicit guard when an operation must wait for clients:
checks:
ftp-control-connections:
type: tcp_connections
port: 21
count: { op: ">", value: 5 }
reports: state
rules:
block-restart-with-active-connections:
type: guard
blocks: [restart, stop]
if: { active: { check: ftp-control-connections } }
then:
action: block
message: "${display_name} has active TCP connections; restart denied"The operation engine evaluates this guard immediately before both manual and automatic actions. An unavailable connection check denies the action rather than treating an observation failure as an empty connection set.
This fail-closed rule applies to every check and to metric:, process: and
changed: condition leaves: a timeout, unreadable source, malformed sample,
missing source or not-ready metric is an unavailable observation, not a valid
false condition. Value and state predicates do not match when an evaluated
leaf is unavailable or skipped, including beneath not, and or or. In particular,
file: { exists: false } requires observed absence, and active never treats
a skipped check as a successful observation. Guards deny the operation on
these gaps. The explicit failed operator retains its check-failure contract: an
unavailable probe counts as a failure for ordinary rules (for example, an HTTP
probe timeout), but a skipped probe does not. Guards still reject unavailable
probes even under failed. A valid sample that simply does not satisfy its
predicate remains an ordinary false result. For process: leaves and type: process
checks, an incomplete process table or a user that does not resolve (for
example an NSS/LDAP outage) proves neither presence nor absence and is
unavailable; a live match is still reported as running.
For MySQL/MariaDB and PostgreSQL, use a read-only sql check that counts the
application sessions you want to preserve; this is more accurate than sockets
and requires a monitoring account allowed to see other sessions. Redis/Valkey
already exposes connected_clients through its redis check, and Memcached
exposes curr_connections. Ready-to-copy service overrides cover
MySQL,
PostgreSQL,
Redis, and
ProFTPD.
The shipped MySQL, MariaDB and PostgreSQL catalog services include a default
optional backup process check and a
block-restart-during-backup guard. The check matches common local backup
tools by exact resolved executable path (exe_any) and database backup user
(backup_user, defaulting to mysql or postgres). Override that check locally
when your backup runs under another user or from non-standard paths. If a
logged-in terminal user runs sermoctl restart while this backup guard blocks
the action, Sermo also sends that user a best-effort native TTY notice; cron and
other non-interactive runs are not notified.
The examples
examples/services/mariadb-backup-guard.yml
and
examples/services/mysql-wal-g-backup-guard.yml
show the same shape for extra app-linked tools or site-specific overrides. The
apps: list is an override, so a service that adds a backup app must keep the
database app too, for example apps: [mysql, wal-g-mysql] or
apps: [mariadb, wal-g-mysql].
For PostgreSQL site-specific WAL-G paths, use the concrete materialized catalog
daemon and app for the installed version (for example postgres-16) and add
wal-g-pg:
name: postgres-main
uses: postgres-16
apps: [postgres-16, wal-g-pg]
checks:
wal-g-pg:
type: process
optional: true
exe_any: ["${wal_g_pg_binary}", /usr/local/bin/wal-g-pg]
user: postgres
state: running
rules:
block-restart-during-wal-g-pg:
type: guard
blocks: [restart, stop]
if:
active:
check: wal-g-pg
then:
action: block
message: "${display_name} WAL-G backup is running"Guards are evaluated before remediation; a remediation action that a guard blocks never runs.
message: strings may use runtime built-ins. ${date} is the current RFC3339
timestamp, ${event} is the firing rule's name and ${action} is the rule's
primary action. ${rule.duration} is the configured rule span (10m,
3 cycles, or current cycle) and ${rule.window} is the fuller window
description (for 10m, within 15 cycles (min 3), immediate). ${service}
and ${host} are resolved during configuration.
For rules whose condition has exactly one direct check or metric leaf, alert
messages may also use ${check.name}, ${check.type}, ${check.metric},
${check.scope}, ${check.op}, ${check.threshold} and ${check.value}.
Complex conditions with multiple checks leave those ${check.*} values empty
instead of guessing which check should describe the alert.
Byte-valued metric placeholders use the same B through TB presentation as
recovery events, for both the current value and the threshold.
For rules driven by a changed: leaf, alert messages may use
${change.path}, ${change.library}, ${change.app}, ${change.level},
${change.old_version} and ${change.new_version}. Version old/new values are
filled for changed: {app: ...} rules.
rules:
alert-if-memory-high:
type: alert
if:
active:
check: memory-high
for:
duration: 10m
then:
action: alert
message: >-
During ${rule.duration}, ${service} ${check.metric} stayed above
${check.threshold} (current ${check.value}) at ${date}A sql check turns a scalar query into the same kind of threshold watch. Its
value is compared numerically and does not accept the K/M/G size
suffixes that storage byte fields take, so a query that reports a size should
convert it in SQL and name the unit in the message. The PostgreSQL catalog
service uses this shape to watch WAL retained by a replication slot:
watches:
alert-if-replication-slot-backlog:
check:
type: sql
engine: postgres
host: 127.0.0.1
port: ${port}
user: ${monitor_user}
database: ${database}
optional: true
query: >-
SELECT round(coalesce(max(pg_wal_lsn_diff(pg_current_wal_lsn(),
restart_lsn)), 0) / 1048576.0, 1) FROM pg_replication_slots
op: ">"
value: 1024 # MiB, a plain number: no size suffix here
for:
duration: 10m
then:
action: alert
message: >-
During ${rule.duration}, ${service} kept retaining ${check.value} MiB of
WAL for a replication slot (limit ${check.threshold} MiB)Because a sql check reports "not ok" when the connection itself fails, a
database that is down never raises the threshold alert — cover that case with a
separate connection check (type: postgres, type: mysql, …) rather than by
relaxing the query.
policy:
cooldown: 5m
max_actions: 5
max_actions_window: 1h
backoff: { initial: 1m, factor: 2, max: 30m }Policy gates automatic remediation (only sermod, never manual sermoctl
actions): an action is suppressed within cooldown (extended by backoff
after consecutive remediations) or once max_actions is reached in the window.
for/within decide when a rule fires; policy decides whether it may act
now.
backoff grows the effective cooldown after each consecutive remediation:
initial the first time, then multiplied by factor each subsequent time,
capped at max. factor defaults to 2 when omitted (or set to ≤0).
Automatic-remediation state is also persisted in paths.state: LastActionAt,
recent action timestamps used by max_actions, and the current backoff survive a
sermod restart, so restarting the daemon does not bypass cooldown or rate
limits.
Use dry_run: true on a service (or in defaults) when you want remediation
rules to evaluate windows, guards and policy without executing the resulting
start/stop/restart/reload/resume operation. It emits dry-run events and does
not advance live remediation cooldown state. Dry-run also suppresses automatic
rule notifications except wall; manual operator actions are unaffected. See
configuration for examples covering watches and
global defaults.