Skip to content

fix(rust_brain): preserve HLC on snapshot restore, bulk_write, and gossip - #78

Draft
cursor[bot] wants to merge 1 commit into
mainfrom
cursor/critical-bug-inspection-5077
Draft

fix(rust_brain): preserve HLC on snapshot restore, bulk_write, and gossip#78
cursor[bot] wants to merge 1 commit into
mainfrom
cursor/critical-bug-inspection-5077

Conversation

@cursor

@cursor cursor Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Bug and impact

After restore_from_file(), nodes were assigned fresh HLC timestamps instead of the values stored in the snapshot. Causal successor writes and gossip replays with pre-crash HLC were rejected with TimestampRegression, causing silent data loss in disaster-recovery and multi-node deployments.

Concrete trigger: Snapshot a node with hlc=(5000, 10, "nodeA"), restore, then write a causally later value with hlc=(5000, 11, "nodeA") — the write is rejected because restore assigned a fresh wall-clock HLC.

Root cause

restore_from_file() constructed MemoryNode without passing the stored hlc field, so the dataclass default (_hlc.now()) was used. The same gap existed in bulk_write() (no HLC passthrough) and gossip.receive() (no HLC handling).

Fix

  • Add _parse_hlc() for wire/snapshot HLC normalisation (with legacy fallback from ts_ns)
  • restore_from_file: restore hlc, call _hlc.update() per node, hold lock during restore
  • bulk_write: pass through hlc from row payloads
  • gossip.receive: apply hlc, reject stale updates, skip missing-HLC overwrites on existing keys

Validation

  • Reproduced the bug locally (successor write raised TimestampRegression)
  • 38 targeted tests pass: test_hlc_snapshot_gossip.py, test_enterprise_backup.py, test_gossip.py, test_rust_brain.py, test_rust_brain_concurrency.py
Open in Web View Automation 

…ssip

After snapshot restore, nodes were assigned fresh HLC timestamps instead of
the values stored in the snapshot. Causal successor writes and gossip
replays with pre-crash HLC were rejected with TimestampRegression,
causing silent data loss in disaster-recovery and multi-node deployments.

- Add _parse_hlc() for wire/snapshot HLC normalisation
- restore_from_file: restore hlc, update global HLC, hold lock during restore
- bulk_write: pass through hlc from row payloads
- gossip.receive: apply hlc, reject stale updates, skip missing-hlc overwrites
- Add regression tests in test_hlc_snapshot_gossip.py

Co-authored-by: Daniel <DJLougen@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant