Skip to content

Forward-port #912, #917 and #911 (0.6.1 IBD and drain fixes) to master - #921

Merged
bkeroack merged 3 commits into
masterfrom
fix/fwdport-061-drain-and-pacer
Oct 6, 2026
Merged

bkeroack merged 3 commits into
masterfrom
fix/fwdport-061-drain-and-pacer

Conversation

@bkeroack

@bkeroack bkeroack commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Brings three fixes merged into release/0.6 for 0.6.1 back to master, one commit each, cherry-picked with -x:

No master commit since the release/0.6 fork touches the files these change, so all three apply unchanged. The code patch is identical to the one on release/0.6: git patch-id of e10c19e1..98bca53e (code paths) equals that of this branch's diff.

Release notes

The three CHANGELOG.md bullets and their docs/release-notes/0.6.1-pre.md sections stay on release/0.6. They describe 0.6.1 and will come to master with the 0.6.1 cut, as #901 brought 0.6.0 and #910 left #899's. Nothing is added to 0.7.0-pre.md.

Merging

"Rebase and merge" keeps the three commits separate, each with its cherry picked from line.

Verified

Ran locally on this branch:

  • cargo clippy --all-targets --all-features --locked -- -D warnings;
  • cargo test --test e2e --no-run --locked --features e2e;
  • the node lib suite (2132 passed);
  • the regtest stored-tail, -blocknotify, -stopatheight and ping/pong/sync tests (11), on the first two commits.

On release/0.6, each PR passed CI and a Devin review before merging. Push CI passed on the tree after #917.

🤖 Generated with Claude Code

bkeroack and others added 3 commits October 6, 2026 15:18
…912)

* Report blocks the stored-tail drain connects as chain events (#900)

The drain connects a stored block once its parent is the tip: a block
that arrived before its parent, or the tail an IBD batch leaves behind.
It connected through connect_stored_block, which emits nothing, so
-blocknotify, block announcement, Electrum and Esplora subscribers, the
streaming API and ZMQ hashblock never heard of that block.

The drain now emits BlockConnected after the connect and after the
mempool has dropped the block's transactions, as accept_block does.
ChainState::emit_chain_event becomes pub(crate) for it;
connect_stored_block still emits nothing, since the IBD connector uses
it too and stays quiet.

The drain keeps its own -stopatheight check: the watcher in main now
hears of drained blocks, but only after the walk has moved on. The
watcher uses send_if_modified, so when both ask for shutdown the second
neither logs nor wakes shutdown waiters again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Report a drained block before releasing the accept lock (#900)

The stored-tail drain emitted BlockConnected after connect_stored_block
had released accept_lock. accept_block reports a connect while it still
holds the lock, so a submitblock of the drained block's child could
connect and report in between, and subscribers heard of the child
before its parent. connect_stored_block_and_report now purges the
mempool and emits under the lock, as accept_block does, and
emit_chain_event is private again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6856ee9)
* p2p: run the manager loop's maintenance on time under load (#909)

The manager loop's maintenance pass holds IBD stall detection: the
release of a height stuck with one peer (15 s at the connect cursor),
the drop of silent peers, and work for idle peers. Its cadences count
passes ("every 4 ticks (2s)"), so it has to run every 500 ms.

Under initial block download it did not. The fast drain added in #781
set drain_only, so every run of fast drains ended in a wait for the next
tick, and the pass after that tick found the queue full again. With the
manager storing each arriving block, an fsync apiece, the queue is full
at every wake. Measured on a fresh mainnet sync over two minutes, the
pass ran 3-4 times instead of ~240, the stale release never fired, and
the block at the connect cursor sat with one slow peer for up to 53 s.

DrainPacer decides after each drain. Maintenance is due on an interval
tick, as before, or once 500 ms have passed since it last started, and
then runs after the drain in progress however full the queue is. A full
queue otherwise goes straight back to draining (#781's throughput), and
does so after maintenance too. A drain_now wake is neither a tick nor
late, so it drains and waits, as #776 requires. The interval uses
MissedTickBehavior::Delay, so the ticks a busy stretch skipped do not
fire as a burst of passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(release-notes): cite only runs of the shipped #909 fix

Six runs of a release build of this branch against nine of the 0.6.0
release binary, replacing figures that mixed in earlier versions of the
fix and an instrumented build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 89ce0e5)
…911)

* IBD: warn that the connector is stuck only after a 10 s wait (#904)

The connector logged "Connector stuck waiting for block data" at WARN
the first moment it waited on any height. Blocks arrive out of order and
the connector takes them in height order, so early in a healthy sync
that was nearly one warning per block: 112 of the first 624 log lines on
a fresh mainnet sync, 490 of 508 in a fresh Umbrel install.

StuckWait tracks which height the connector waits on and since when.
The warning fires once that wait reaches 10 s, then every 60 s while it
lasts, and carries waited_secs. A new height, or a successful connect,
starts over. In the #904 run the waits worth a warning had the block in
flight 23-60 s on one peer; 97 of the 112 warnings were under 5 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* IBD: raise the stuck-wait warning to 20 s, above the stale-block release (#904)

The scheduler takes a block at the connect cursor back from a slow peer
after 15 s and asks another. A 10 s threshold warned during waits that
release was about to end; 20 s warns only once it has not.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 98bca53)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment