Skip to content

Latest commit

 

History

History
502 lines (415 loc) · 25.4 KB

File metadata and controls

502 lines (415 loc) · 25.4 KB

CLI

sermoctl is the operator and scripting interface. Run it with no arguments or --help for the command index, and use sermoctl help COMMAND or sermoctl COMMAND --help for focused usage, flags and examples.

Root flags

--config /etc/sermo/sermo.yml
--backend auto|systemd|openrc
--json
--quiet / -q
--timeout duration
--version / -V
--help / -h

Global flags may be placed before or after the command. Command-specific flags are shown by sermoctl help COMMAND.

Without --timeout, live service queries (status and is-active) use the 10-second engine check budget; service operations (and watch pause|resume) use engine.operation_timeout (default 90s), the same budget as the daemon and the Web UI, which a service's stop_policy may raise. sessions uses the same 90-second default (a statement kill runs through the daemon's operation engine), while each request to the daemon keeps its 10-second web client bound. notifier test uses engine.default_timeout (default 10s), like the Web UI's test button. Other short probe commands keep their 2-second CLI budget. --timeout must be a positive duration; 0 or a negative value is a usage error (exit 64). For service operations, an explicit --timeout bounds backend preparation and every operation phase after configuration and audit-store initialization. Stop-policy budgets cannot extend this deadline; expiration prevents subsequent lifecycle steps. Recording the outcome retains its separate storage timeout.

sermod daemon flags

sermod is the long-running monitoring daemon. Packaged units normally start it with the standard config path:

sermod run --config /etc/sermo/sermo.yml

Manual runs support these flags:

sermod run [--config PATH] [--verbose|-v]
sermod version
sermod --version
  • --config PATH loads the global config file. The default is /etc/sermo/sermo.yml. Use the same path with sermoctl --config when validating or reloading a non-standard tree.
  • --verbose / -v enables debug logging, including config load details, backend detection and monitor-target counts.

Use sermoctl daemon reload to ask a running daemon to re-read the config file it was started with.

Command surface

sermoctl help [COMMAND]
sermoctl backend
sermoctl version
sermoctl status SERVICE
sermoctl is-active SERVICE
sermoctl watch status WATCH
sermoctl watch monitor WATCH
sermoctl watch unmonitor WATCH
sermoctl watch probe WATCH
sermoctl watch pause RAID_WATCH --confirm MD_ARRAY
sermoctl watch resume RAID_WATCH
sermoctl start SERVICE [--no-cascade]
sermoctl stop SERVICE [--no-cascade]
sermoctl restart SERVICE [--no-cascade]
sermoctl pause SERVICE                  # libvirt VM (suspend) or Docker container only
sermoctl resume SERVICE
sermoctl reload SERVICE

sermoctl mount TARGET                 # TARGET is a configured mount name or absolute path
sermoctl umount TARGET
sermoctl mount status TARGET
sermoctl mount list

sermoctl preflight SERVICE
sermoctl processes SERVICE
sermoctl reap SERVICE [--apply]        # list the service's stray processes; --apply signals the authorized ones
sermoctl sessions [list] [SERVICE]     # SSH, tmux/screen sessions and running database statements
sermoctl sessions kill SERVICE WATCH ID [--connection]   # cancel one listed statement (or close its connection)
sermoctl locks SERVICE
sermoctl monitor SERVICE
sermoctl unmonitor SERVICE

sermoctl panic on|off|status          # daemon-wide emergency switch (see Panic mode)

sermoctl config validate

sermoctl web hash-password [--stdin|--generate] [--hash bcrypt|sha256] [--cost N] [--name LABEL]
                                       # print one credential line for web.password_file

sermoctl daemon reload                 # reload sermod config, not services
sermoctl notifier test NAME            # send an explicit test message through one notifier

sermoctl services [--notify NAME[,NAME]|all]                         # configured services (init, docker, VMs)
sermoctl services catalog [all] [--long] [--notify NAME[,NAME]|all]   # catalog inventory
sermoctl apps [all] [--long]                                          # catalog apps (see Catalog inventory)
sermoctl libs [all] [--long]
sermoctl patterns [catalog]                                           # pattern sets in use (catalog: all)

sermoctl sla [TARGET]                   # availability windows for every service and availability watch, or one
sermoctl sla --series TARGET [--since DURATION]   # per-minute series; --since default 24h

sermoctl events [SERVICE] [--limit N]   # list recent events (global or for SERVICE); SEVERITY shows a graded event's level
sermoctl events clear [--before TIME]   # omit TIME to clear all; TIME may be non-future RFC3339 or positive duration
                                        # only events strictly before the timestamp are removed
sermoctl activity clear [--before TIME] # clears the same log shown in Events

sermoctl state compact [--before TIME]  # consolidates and prunes stored history, then vacuums the state database
                                        # omit TIME for the configured retention; TIME additionally drops older history (non-future RFC3339 or positive duration)

sermoctl lock SERVICE [--name NAME] --reason REASON --ttl DURATION -- COMMAND...
sermoctl lock acquire SERVICE [--name NAME] --reason REASON --ttl DURATION
sermoctl lock release SERVICE [--name NAME]

sermoctl wizard
sermoctl wizard service|docker|vm|mount|volume|net|uplink

Availability

sermoctl sla is observed check availability: it only counts monitored daemon cycles. A window with no observed cycles reads n/a, not downtime — daemon downtime or missing data never becomes observed downtime. sermoctl sla --series TARGET emits that target's stored per-minute availability series (the raw data a graph is built from).

A target is a configured service or a host watch whose check asserts availability — tcp, ports, http, route, the state metric of net and icmp, and the endpoint form of cert. Those are the checks whose failing half is genuinely something not answering; a cert check reading a file on disk is a certificate nearing expiry, which is a condition, not a host being unreachable. A condition watch keeps no series: a filesystem crossing 90% used is a threshold being met, not an outage, and reporting it as availability would give a percentage that reads like uptime while meaning something else. The same exclusions the services already apply carry over — a verdictless watch (reports: state) is a sensor and an advisory (severity: warning) is a thing to look at, so neither is downtime. A name is resolved as a service first, so an existing service name never changes meaning.

Examples:

sermoctl help restart
sermoctl restart mysql-main
sermoctl services --notify ops-email
sermoctl notifier test ops-email
sermoctl daemon reload
sermoctl state compact --before 720h

Panic mode

Panic mode is a daemon-wide emergency switch for maintenance windows, attacks, denial-of-service, system malfunction or overload. While it is on, the daemon keeps running its checks (so status stays visible) but suspends all hooks, alert notifications and automatic remediation. Manual operations (start, stop, restart, reload, pause, resume) stay available, so you can drive services by hand without the daemon fighting you.

sermoctl panic on        # suspend hooks, alerts and automatic remediation
sermoctl panic status    # show the current state (default when no argument)
sermoctl panic off       # resume normal operation

The flag is persisted in the state database (paths.state), so it survives daemon restarts until you turn it off, and the CLI works without the web UI enabled. The running daemon picks up a change within ~1 second. While active, the daemon status reported by /readyz and the web header shows panic mode. In the web UI the same toggle is the red panic mode button in the footer (it asks for confirmation in both directions so it is not triggered by accident). The CLI applies the change immediately without a prompt.

If the initial state read fails after daemon startup, automatic side effects stay suspended until a successful read establishes the panic flag. Checks and manual operations remain available. Later read failures retain the last successfully read value. Failed reads are retried after the same one-second cache interval, measured from the end of the previous read.

Service target resolution

For a configured service, sermoctl status, is-active and service operations resolve the same control target that sermod and the web UI use. When sermod is running with web enabled, sermoctl status prefers the daemon's computed state (including starting during startup settling); if the web API is unreachable it falls back to the init backend plus local monitor metadata, as before. Service states are: disabled, stopped, started (backend active but not monitored), starting (startup/operation settling), collecting (active and monitored, but graphs/indicators are not complete yet), warning (active with an advisory problem such as an invalid application configuration, an unattributed process tree, or an init unit failed while an exact process and its functional checks remain healthy), restart_required (active and observed but running a binary that was replaced on disk), monitored (active, monitored and observability-ready) and failed. Without the daemon view, a configured active monitored service falls back to collecting; an active service that is not known to be monitored falls back to started.

For the second warning case, status remains failed in the API: operations still follow the init backend and retain their normal locks, guards and preflight gates. The warning only prevents a working workload from being presented as an application outage.

sermoctl status SERVICE exposes warning and restart_required directly. For a configuration warning, sermoctl preflight SERVICE reruns the same bounded preflight.config command and prints its current output. is-active continues to report only the init backend's active/inactive verdict.

A backend status of unknown is not a verdict of "down" — a transitional systemd state such as activating/deactivating, an init script that replaces status with its own report, or a query that timed out can read unknown while the service runs normally — so it never yields failed on its own. The service's own checks decide instead: a failing required check still reads failed, and healthy checks read active or collecting rather than monitored, because a backend that would not answer cannot underwrite the full-observability claim.

Each manual service operation persists exactly one result in the shared event feed, so sermoctl events and the Web UI show the same action outcome. Sermoctl does not start an operation when the state database cannot be opened for that audit record.

sermoctl is-active is different: it always probes the init backend (active / inactive / paused) for the exit code and plain-text output. A monitored service still settling with an inactive backend therefore shows state=starting in status but exits 1 from is-active until the unit reports active.

The same preference applies to the STATUS column of sermoctl apps for installed applications monitored by the daemon. Catalog apps whose binary is not installed are omitted from sermoctl apps and do not participate in startup settling.

Only the daemon observes watches, so sermoctl watch status WATCH reports the daemon's computed state. WATCH must be a configured watch (a host watch or "<service>:<watch>"); an unknown name is an error. When sermod is stopped or its web API does not answer (web disabled, token rejected), the command prints state=unknown (also in --json), warns on stderr and exits 2, so a monitoring script never reads an unobserved watch as healthy.

When the daemon has current watch readings, sermoctl watch status WATCH also prints them (including RAID operation and rebuild percentage) and the separate last-check timestamp; --json exposes the same readings in a readings array.

sermoctl watch monitor|unmonitor WATCH pauses or resumes a single watch, persisted under paths.state and read live by the daemon. WATCH is a host watch name or a service-embedded watch "<service>:<watch>"; a watch's monitor state is independent of its service's, so unmonitor on a service never pauses its watches. Repeating either command after that state is already effective is a successful no-op and preserves the original monitor source and change time.

sermoctl watch probe WATCH asks the running daemon to run one fresh sample for a host diskio, hdparm, lvm, raid, smart, storcli or ssacli watch and prints the resulting readings when available (for hardware RAID this includes one reading per controller, cache, virtual volume and physical drive, with identity/capacity, controller RAM, controller SMART and rebuild progress where reported). Every probe except smart is read-only. A smart probe starts the device's short SMART self-test with smartctl --test=short DEVICE; success means the device accepted the test, not that it has passed it. Normal scheduled SMART checks remain read-only health/attribute reads. The command records a probe event and last-check time, but does not run rules, notifications or remediation. Its status line names the result's severity: OK, or DEBUG, INFO, WARN, FAIL (error) or CRIT for a failing sample. A RAID watch with raid_control.pause_resume: true and an explicit check.array also supports watch pause and watch resume. Pausing requires --confirm MD_ARRAY in addition to naming the watch; both actions re-check the array, use an exclusive runtime operation lock and verify the resulting kernel state. Pause freezes the array (sync_action frozen); resume accepts any currently frozen configured array, including one frozen outside Sermo.

The daemon records both probe/running when a manual sample starts and its probe/ok or probe/failed completion event with the elapsed time. A SMART self-test remains testing in watch status until the device reports it has ended. RAID/LVM device work is also reported as testing, recovering, rebuilding, repairing, moving or merging, including the reported percentage where available; those states describe work, not health. Only one manual sample for a watch may run at once; sermoctl watch probe waits for that same daemon task and reports an already-running sample instead of starting a second disk, LVM, RAID or SMART command. Sermo reads the service's service: candidates, picks the first unit known by the active backend, and normalizes systemd names with .service when needed.

If the backend probe cannot surface a configured init unit but the service still has a usable configured seed, Sermo falls back to that unit and prints a warning, matching the daemon/web behavior used for historic init-service setups. There is no fallback for invalid control: targets or a per-backend service: map with no candidate for the active backend; those are configuration errors.

Configured services

sermoctl services lists every service configured under paths.services, whatever its control backend: init units, Docker containers (control: {type: docker}) and libvirt VMs or networks (control: {type: libvirt|libvirt-network}). Disabled services are listed as disabled.

SERVICE      TYPE     STATE      MONITORED
nginx-main   systemd  monitored  yes
web-ctr      docker   paused     yes
vm-web01     libvirt  disabled   no

TYPE is the control backend. When sermod answers on its web API, STATE and MONITORED are the daemon's computed view (one GET /api/services), the same as the web UI Services panel. A service the daemon does not report, such as one added since its last reload, or every service when sermod is down, is probed locally the way sermoctl status does. A service that does not resolve is listed with state error and a warning on stderr.

--json prints {"services": [{"name", "display_name", "backend", "unit", "state", "monitored", "enabled", "error"}]}. --notify sends a configured services health report: a monitored service in a failed, warning, stale, restart-required, stopped or error state counts as an issue, and disabled or unmonitored services are counted apart.

sermoctl patterns lists only the output-analysis pattern sets that configured services name in analyze.use, with their rule count and the services using them (USED BY).

Catalog inventory

sermoctl services catalog, sermoctl apps, sermoctl libs and sermoctl patterns catalog list catalog definitions shipped in the packaged catalog (see services.md): which profiles are installed, the version their version command reports, and whether they resolve. Add all to include entries whose binary or library file is not present on the host.

Question Where to look
Which services does this host supervise, and in what state? sermoctl services
Which catalog service profiles exist / are installed? sermoctl services catalog [all]
Which catalog apps / libs exist? sermoctl apps, sermoctl libs
Which pattern sets are in use / exist? sermoctl patterns, sermoctl patterns catalog
One configured service's live state sermoctl status SERVICE, sermoctl is-active SERVICE
Availability history for services and availability watches sermoctl sla [TARGET]

The web UI uses the same split: Services shows configured runtime services, like sermoctl services; Applications (GET /api/applications) and Libraries (GET /api/libraries) are installed catalog inventories, aligned with sermoctl apps and sermoctl libs.

Reaping stray processes

A stray is a process the init backend attributes to the service's control group that no processes: selector or pidfile claims and that no longer hangs off the unit's principal process — a probe that daemonized, a child the daemon never reaped, a survivor of an earlier incarnation. sermoctl processes SERVICE flags them with stray=true, and the injected strays check reports them every cycle.

A stop never reaps, so a restart that a stray blocks ends in orphan_processes and names it; clearing it is this command's job.

sermoctl reap SERVICE is a preview: it lists every stray, reports how many would be signalled, and touches nothing — no operation lock, no event.

sermoctl reap SERVICE --apply signals them, gated by the service's own reap.kill_only_if selector. With no such block nothing is authorized, so the command reports every stray, signals none and exits 75. Otherwise the exit code follows the usual operation mapping: 0 when no stray remains, 1 when one was spared or outlived SIGKILL (orphan_processes), 2 on a failure.

--apply is rejected by every other command, and no rule action can reap — see safety.md for the whole contract and services.md for the reap: block.

sermoctl config validate prints each validation error as ERROR and exits 78. It also prints advisory findings as WARN — today a levels: tier Sermo ignores because it is not stricter than the threshold below it — and still exits 0, since the configuration loads and runs; --json lists them under warnings.

Sessions

sermoctl sessions prints the running daemon's session inventory — the same data as the Web UI's Sessions panel: SSH terminals, tmux/screen sessions and the statements each db_queries watch last sampled. It needs sermod with web.port set (the same credentials as other daemon queries) and never connects to a database itself.

sermoctl sessions                 # every source on this host
sermoctl sessions mysql           # only the rows and sources of one service
sermoctl --json sessions mysql    # the raw inventory (GET /api/sessions), filtered
Database statements:
SERVICE  WATCH                         ENGINE   ID    USER  DB    RUNNING  LONG  QUERY
mysql    alert-if-query-long-running   mariadb  4242  app   shop  1h 2m 3s LONG  SELECT * FROM orders WHERE note LIKE '%gift%' AND created …
mysql    alert-if-query-long-running   mariadb  4250  app   shop  4s       -     UPDATE carts SET touched = NOW() WHERE id = 991

LONG marks a statement past the watch's min_duration and filters (the same rule its alert uses), STOPPING one the server is already stopping; the query column is cut to fit a terminal (use --json for the full, already redacted text). A source that is still collecting or unavailable is reported on stderr with its reason, so an empty table is never mistaken for an idle server.

sermoctl sessions kill SERVICE WATCH ID cancels one statement. It reads the inventory first to obtain the statement's exact identity — an id that is not listed fails with "not listed; refresh" — then asks the daemon to cancel it. --connection closes the statement's connection instead. The daemon re-lists the server and refuses if the statement changed or ended, and the kill goes through the service's operation lock, guards (blocks: [kill_query]) and audit event; see safety.

sermoctl sessions kill mysql alert-if-query-long-running 4242
sermoctl sessions kill postgres alert-if-query-long-running 31337 --connection

Exit codes: 0 cancelled, 75 refused by the daemon (statement changed or finished, guard, lock), 2 when the statement is not listed or cannot be killed, or the daemon cannot be reached. --connection is rejected by every other command.

Exit codes

0   success / active / allowed
1   expected false condition, such as inactive or a failed check
2   internal or runtime error / backend not detected
64  usage error (bad flags or arguments)
75  temporarily blocked action, such as an active backup lock or guard
78  configuration invalid (syntax, schema or `config validate` failure)

The 2 vs 78 distinction: use 78 whenever the problem is in the config files the operator can fix (YAML syntax, missing kind/name, unknown variable, unresolved uses/clone, failed config validate). 2 is everything else that is not a clean false (1), a usage error (64) or a temporary block (75): I/O errors, backend not detected, an exec that could not be launched, an unexpected panic recovered at the top level.

is-active maps directly: 0 active, 1 not active (including a backend-paused container or VM), 2 error. Pausing daemon monitoring with unmonitor does not change the answer.

status and is-active still answer for a raw init unit when there is no config. When the config file exists but does not load (for example a YAML syntax error), they print warning: config not loaded, treating SERVICE as an init unit: … on stderr (unless --quiet), because a configured alias such as mysql-main then reads as an unknown unit.

Named locks

lock acquire and the lock SERVICE -- COMMAND wrapper require a configured service (a name or alias status accepts): the engine checks locks by the canonical service name, so a lock on a mistyped name would protect nothing. lock release also accepts a service that is no longer configured, to clean up its leftover lock; when no such lock exists (for example a mistyped --name) it prints no named lock SERVICE[.NAME] to release and exits 1. With --json, lock acquire prints {"ok":true,"service":…,"lock":…,"path":…} and lock release prints {"ok":true|false,"service":…,"lock":…} (plus message when nothing was released); locks --json and mount list --json print an empty array rather than null when there is nothing to list.

sermoctl lock SERVICE --reason R --ttl D -- COMMAND... holds the named lock while COMMAND runs; the lock's owner is the sermoctl process. sermoctl therefore stays alive until COMMAND exits: it forwards SIGTERM and SIGHUP to COMMAND, and it neither dies on nor forwards SIGINT/SIGQUIT (typed at a terminal they already reach COMMAND through the foreground process group). The lock is released only after COMMAND has exited, so a backup that cleans up after SIGTERM stays protected until it finishes. SIGKILL of sermoctl cannot be caught; the lock then turns stale (dead owner) and stops blocking.

The wrapper exits with COMMAND's exit status, or 128 + signal when a signal killed COMMAND (for example 143 for SIGTERM), as a shell reports it.

Mounts

Mount actions are fstab-backed and use storage watch files with a mount: block from directories listed in paths.watches (the wizard writes /etc/sermo/mounts by default). A path target that is not configured is still accepted, but it uses safe defaults and must exist in /etc/fstab. See storage and mount units. sermoctl umount / is always rejected; Sermo never unmounts the root filesystem. sermoctl umount TARGET --force permits umount -f after the normal unmount fails, --lazy permits umount -l as the last fallback, and --kill-blockers signals only blockers that match mount.stop_policy.kill_only_if.

sermoctl wizard mount lists mount points declared in /etc/fstab and writes safe storage watch files under mounts/, adding that directory to paths.watches; it does not execute mount or umount while generating the config.