Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,14 @@
# Build output of the standalone BDK canary crate (scripts/canary/bdk-canary).
# Its Cargo.lock IS committed (reproducible downstream pin); its target/ is not.
/scripts/canary/bdk-canary/target

# Same for the reference push relay (contrib/push-relay): a standalone crate
# excluded from the workspace, so the root /target rule does not cover it. The
# rule lives this early in the stack, ahead of the crate it protects, because
# 831 MiB of its build output was once committed by a stray `git add -A` four
# PRs before the crate itself existed — and .gitignore does not untrack what is
# already tracked.
/contrib/push-relay/target
*.swp
.idea/
.vscode/
Expand Down
23 changes: 23 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,29 @@ layout) per [`STABILITY_POLICY.md`](STABILITY_POLICY.md).
capability gates on the RPC surface. Not replayable (no cursor —
detectors re-raise standing conditions after a restart). Wire schema only in
this change; the detectors that emit them land next.
- Alerting: node-health detectors. satd now watches six conditions about itself
— stalled tip, low disk, congested mempool, peer starvation, IBD completion,
deep reorg — and reports each through three surfaces at once: a `status`
streaming event, an entry in `getwarnings` (which fires the Core-compatible
`alertnotify` hook), and a `satd_alert_active{kind}` gauge. Standing
conditions raise once and clear once, with hysteresis so a value sitting on
the threshold does not flap. The one-shot events (`ibd_complete`,
`deep_reorg`) fire `alertnotify` and the streaming event but deliberately do
**not** enter `getwarnings` — nothing would ever clear them, and on chains
where multi-block reorgs are routine that would wedge
`getblockchaininfo.warnings` and the TUI modal permanently. Five new hot-reloadable thresholds
(`alerttipstallseconds=3600` — `0` on regtest, `alertdiskfreemb=10240`,
`alertmempoolfullpct=90`, `alertpeerfloor=3` — `0` on regtest, and capped by
the number of `-connect=` peers when that is set —
`alertreorgdepth=3` — `10` on test networks, `0` on regtest); `0` disables a
detector. `deep_reorg` depth, fork
height and new tip are read from the durable reorg log, so they are exact
regardless of chain-event lag. Two new gauges close longstanding observability gaps
independently of alerting: `satd_tip_last_connect_age_seconds` and
`satd_disk_free_bytes`, the latter sampled even when `alertdiskfreemb=0`. The
free-space probe is bounded and carried across polls, so a `blocksdir` on an
unresponsive network mount stalls only the disk alert — never the other
detectors — and strands at most one blocking thread rather than one per poll.

## Releases

Expand Down
27 changes: 27 additions & 0 deletions docs/api/streaming.md
Original file line number Diff line number Diff line change
Expand Up @@ -973,6 +973,33 @@ re-raise the ones still standing, which is what makes health alerting
at-least-once across a restart. A condition that both raised and fully cleared
while a consumer was away is stale by definition and is not reconstructed.

**`details` keys by kind.** Values are strings. Most are decimal numbers, but
not all — parse per key rather than assuming the whole map is numeric.

| Kind | Keys |
|---|---|
| `ibd_complete` | `height` |
| `tip_stall` | raise: `seconds_since_block`, `threshold_seconds`, `tip_height`. Cleared by a block: `height`. Cleared by the poll: `seconds_since_block`, `threshold_seconds` |
| `disk_low` | `free_bytes`, `threshold_bytes`; `clear_threshold_bytes` when cleared by recovered space |
| `mempool_congested` | `bytes_used`, `bytes_cap`, `threshold_pct`, `mempoolminfee_sat_per_kvb` (raise only) |
| `peer_floor` | `peers`, `peers_outbound`, `peers_inbound`, `threshold` (raise only) |
| `deep_reorg` | `depth`, `from_height`, `to_height`, `fork_height` |

`deep_reorg` figures are exact. Depth, fork height and the reconnected chain
are read from the reorg log record that `perform_reorg` writes and fsyncs, not
reconstructed by counting disconnect events off the event bus — so a reorg deep
enough to overrun the bus ring is reported with the same precision as a shallow
one. `to_height` is the new tip, not the first reconnected block.

Any kind may additionally carry `reason` — a non-numeric token on a `cleared`
event emitted because the detector can no longer evaluate the condition, rather
than because the condition recovered:

| Token | Meaning |
|---|---|
| `detector_disabled` | The operator set this detector's threshold to `0`. |
| `mempool_cap_zero` | `maxmempool` is `0`, so there is no occupancy ratio to evaluate. `alertmempoolfullpct` is still armed — this is not the operator turning the detector off, and a consumer that suppresses `detector_disabled` should not suppress this. |

**Additive by construction.** `details` is a string map and `StatusKind` is an
open enum: new kinds and new detail keys ship without a schema bump (§4). A
consumer must tolerate an unrecognized `kind` — `message` and `severity` remain
Expand Down
20 changes: 20 additions & 0 deletions docs/manual/src/config-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -394,6 +394,26 @@ Core ZMQ wire-format compatible.)
| `reorgwebhook` | none | hot | satd | HTTP(S) endpoint receiving a POST on reorg detection. |
| `reorgwebhooksecret` | none | hot | satd | HMAC-SHA256 secret signing webhook bodies via `X-Satd-Signature`. |

## Health alerts

Thresholds for the node-health detectors. Each raises a `status` event on the
[Streaming Consumption API](streaming.md) and an entry in `getwarnings` (which
also fires `alertnotify`) when its condition is entered, and retracts both when
it recovers. Every one is hot-reloadable — retuning an alert should not need a
restart, since you are usually retuning it *because* it is firing.

Set a threshold to `0` to disable that detector. See
[Observability → Node-health alerts](observability.md#node-health-alerts) for
the taxonomy and the details each event carries.

| Key | Default | Reload | Compat | Description |
|---|---|---|---|---|
| `alerttipstallseconds` | `3600` (`0` on regtest) | hot | satd | Raise `tip_stall` after this many seconds with no connected block. Defaults to disabled on regtest only, where blocks exist just when a test mines them and an idle chain is normal; every other network — test networks included — keeps the hour, since going an hour without a block is not an ordinary property of thin hashrate the way a shallow reorg is. *Not* suppressed during initial block download — `is_initial_block_download()` compares the tip header's timestamp against the wall clock rather than tracking sync progress, so a node that was caught up and then wedged re-enters it precisely when you need paging. A node that is genuinely syncing connects blocks continuously and so never crosses the threshold. Cleared the moment a block connects — or, if you raise this value past the current tip age, on the next detector poll. |
| `alertdiskfreemb` | `10240` | hot | satd | Raise `disk_low` below this many MiB free on the blocks directory (or the data directory when `blocksdir` is not split out). Clears at 1.5× the floor, or as soon as you lower the floor below the current reading. |
| `alertmempoolfullpct` | `90` | hot | satd | Raise `mempool_congested` at this percentage of `maxmempool`. Clears below 75 % of the raise line, or as soon as you raise the threshold above the current occupancy. Values above 100 are clamped. |
| `alertpeerfloor` | `3` (`0` on regtest; capped by the `-connect=` count) | hot | satd | Raise `peer_floor` below this many connected peers (inbound + outbound). The count must hold for 60 s in either direction, so ordinary peer churn does not page you, and it does not raise until 90 s after startup or the first peer, whichever is sooner. Defaults to disabled on regtest only, where a node with no peers is normal; **signet keeps the floor** — set `alertpeerfloor=0` explicitly on a deliberately isolated signet node. When `connect=` is set the default drops to that many peers (never above `3`), since `connect=` suppresses DNS and fixed seeds and the node can never exceed the addresses you named — an explicit value here still overrides it. |
| `alertreorgdepth` | `3` (`10` on test networks, `0` on regtest) | hot | satd | Emit the one-shot `deep_reorg` event for a reorg that rolls back at least this many blocks. Depth 3 is an incident on mainnet and ordinary on a chain with thin, volatile hashrate, so signet/testnet/testnet4 default to `10` — above the 6-confirmation convention, so a reorg that invalidated something a wallet called final still reports. Regtest defaults off: its test suites reorg deliberately. |

> **Note.** The `*notify` shell hooks (`blocknotify`, `alertnotify`,
> `startupnotify`, `shutdownnotify`) exist for drop-in Bitcoin Core
> compatibility and quick scripts. They are best-effort shell execs with no
Expand Down
100 changes: 100 additions & 0 deletions docs/manual/src/observability.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,106 @@ labels) and does not consume an RPC worker on every scrape.
> one-time handshake bytes are not included, so absolute socket totals read
> marginally lower than the kernel's.

## Node-health alerts

Metrics tell you what the node is doing; health alerts tell you when it has
stopped doing it. satd watches six conditions about *itself* and reports each
one through three surfaces at once, so they can never disagree:

* a `status` event on the [Streaming Consumption API](streaming.md) (category
bit 16 — see §7.8 of the wire spec),
* an entry in `getwarnings` (and therefore in `getblockchaininfo.warnings`
and the TUI), which also fires the Core-compatible `alertnotify` hook,
* a `satd_alert_active{kind="..."}` gauge on `/metrics`.

**One-shot events are the exception to the middle surface.** `ibd_complete` and
`deep_reorg` describe something that *happened*; there is no state for anything
to later clear. They fire `alertnotify` and emit their `status` event, but they
do **not** create a `getwarnings` entry. An entry nothing clears would sit in
`getblockchaininfo.warnings` for the life of the process and hold the TUI's
warning modal open — which on signet and testnet4, where reorgs several blocks
deep are ordinary, would happen on the first one and never stop. The durable
record of a reorg is the reorg log (`getreorghistory`), not the warnings set.

| Condition | Severity | Raises when | Clears when |
|---|---|---|---|
| `ibd_complete` | info | initial block download finishes | one-shot |
| `tip_stall` | critical | no block connected for `alerttipstallseconds`, outside IBD | the next block connects, or the threshold no longer considers the tip stalled |
| `disk_low` | critical | free space below `alertdiskfreemb` | free space reaches 1.5× the floor, or the floor is lowered below the current reading |
| `mempool_congested` | warning | mempool at `alertmempoolfullpct` of its cap | occupancy drops below 75 % of the raise line, or the threshold is raised above the current occupancy |
| `peer_floor` | warning | fewer than `alertpeerfloor` peers for 60 s (after a 90 s startup grace) | at or above the floor for 60 s |
| `deep_reorg` | critical | a reorg rolled back ≥ `alertreorgdepth` blocks (default `3` on mainnet, `10` on test networks, off on regtest) | one-shot |

Every standing condition raises **once** on entry and clears **once** on
recovery — you get a pair of events, not a stream of repeats — and the gap
between the raise and clear lines (a ratio, a hold time, or both) means a value
sitting on the threshold does not flap your pager. `ibd_complete` and
`deep_reorg` describe things that happened rather than states that persist, so
they are one-shot: they never clear, and for the same reason they never enter
`getwarnings` at all.

Thresholds are configured with the `alert*` keys in the
[Configuration Reference](config-reference.md#health-alerts); all of them are
hot-reloadable, and setting one to `0` disables that detector. Each event
carries a `details` map with the numbers behind it (free bytes and the floor,
seconds since the last block and the tip height, the reorg's true depth and
fork height, the mempool's current `mempoolminfee`), so an alert is actionable
without a follow-up query. The watched *path* is deliberately not in the event —
it goes to every `status` subscriber and into push-notification bodies, and an
absolute datadir path usually names the account it runs under. The node logs it
instead.

**Retuning a threshold always clears its own alert.** The gap between each raise
and clear line stops a value hovering at the threshold from flapping, but it
would otherwise trap the operator who raises a threshold *because* the alert is
firing: the unchanged reading lands between the new raise line and the new clear
line, where neither fires. So a standing condition also clears when the
threshold moves such that it would no longer raise. Without this,
`mempool_congested` in particular was inescapable — `alertmempoolfullpct` clamps
at 100 and the clear line is 75 % of the raise line, so past 75 % occupancy no
setting could clear it.

`alertreorgdepth` defaults to `3` on mainnet, where a reorg that deep costs real
hashrate and invalidates transactions merchants have started treating as
settled. Signet, testnet and testnet4 default to `10`: those chains are not
economically secured, and reorgs a few blocks deep are an ordinary consequence
of thin, volatile hashrate rather than an incident. Defaulting them to mainnet's
sensitivity would run `-alertnotify` for the network working as designed, and an
alert that fires during normal operation is one you learn to ignore — which
costs you the mainnet alert too. It is raised rather than switched off because
past the 6-confirmation convention a wallet has been told something false, and
that is worth reporting on any chain. Regtest is off entirely; its test suites
reorg deliberately. Set `alertreorgdepth=3` explicitly if you want mainnet
sensitivity on a test network.

`alertpeerfloor` defaults to `3` everywhere except regtest, where it is `0`
(disabled) because running with no peers at all is a regtest node's normal
operating state rather than a fault. Signet keeps the floor: it is a public
network with real peers, and a detector defaulted off is indistinguishable from
a healthy one — `satd_alert_active{kind="peer_floor"}` reads `0` either way. Set
`alertpeerfloor=0` explicitly on a deliberately isolated signet node.

Where the floor is active, a node that has never seen a peer gets a 90 s startup
grace, and the ordinary hold begins when that grace expires or when the first
peer arrives, whichever comes first. The grace defers the start of the hold
rather than shortening it, so a node still dialing out does not page anyone on
the way up.

**Durability.** Health events are not replayable: they carry no resume cursor,
and a `from_cursor` reconnect never yields one. Instead the detectors
re-evaluate from scratch on startup and re-raise anything still standing, so a
consumer that was disconnected across a restart still learns about a live
problem. A condition that both raised and cleared while nothing was listening is
stale by definition and is not reconstructed. For the same reason, a subscriber
that attaches *after* a condition raised will not see it until the condition
changes — check `getwarnings` for current state on connect.

Two of the gauges are useful independently of alerting:
`satd_tip_last_connect_age_seconds` (seconds since the last connected block)
and `satd_disk_free_bytes` (free space on the watched directory). The latter is
omitted rather than reported as zero when the filesystem cannot be
interrogated.

## Structured JSON Logging

`satd` logs to stdout. Use `--log-format=json` to switch from the text format
Expand Down
Loading
Loading