fix: self-heal manifest-unreferenced branch forks (stop wedged branches) - #231
Merged
Conversation
The global Arc<RwLock<Omnigraph>> that once serialized every server write was removed — the server holds the engine as a lockless Arc<Omnigraph> and write methods are &self, so the per-(table_key, branch) write queues are now the actual write-serialization mechanism (in-process only). Correct comments that still claimed the global lock is 'still in place' / 'today', or framed the queues as MR-686 scaffolding: write_queue.rs module doc, exec/merge.rs, db/omnigraph/schema_apply.rs, db/manifest/recovery.rs, and the bench_concurrent_http.rs example (which also wrongly stated mutate_as is &mut self). workload.rs is left as-is — its 'previous global RwLock' wording is accurate history.
An interrupted first-write fork (create_branch succeeded, the manifest publish did not) leaves a fully-formed Lance branch ref the manifest never references. The branch stays a valid manifest branch, so cleanup's reconciler never reclaims it, and today the next write to that table wedges with 'incomplete prior delete; run cleanup'. Forge that exact residue (a live 'feature' branch + a directly-created 'feature' ref on the Person table the manifest doesn't reference) and assert the next load AND mutate self-heal. Deterministic and local — no S3 or timing, since the forge IS the post-crash state. Adds a shared node_table_uri helper. This commit is RED: it reproduces the bug and fails against the unfixed engine with the predicted symptom. The fix follows in the next commit.
The first write to a table on a branch lazily forks it via Lance create_branch, a durable two-phase op that advances Lance state BEFORE the atomic manifest publish. If the writer dies or its request future is cancelled between the fork and the publish, the branch ref is fully formed but the manifest never references it. The next write re-enters the fork path, create_branch collides, and the engine wedged with 'orphaned table state ... incomplete prior delete; run cleanup' — which cleanup could not even fix, because the branch is still a live manifest branch. This hit load, mutate, ingest, and the merge fork path (one shared engine chokepoint), so a routine deploy restart or client disconnect could wedge a branch. Fix: treat the per-table fork ref as derived state of the manifest. fork_branch_ from_state returns a typed ForkOutcome instead of a human 'incomplete prior delete' error; on RefAlreadyExists the db layer reclaims the manifest- unreferenced fork (force_delete_branch + re-fork, exactly once) and proceeds. A live committed fork is still routed to a retryable conflict before the fork path, so concurrent first-writes stay correct. Reclaim is only safe if no in-process writer can be mid-fork, so the write entry points (load, mutate) acquire the per-(table, branch) write queues for all touched tables up front — before the fork, held through the publish — when forking a non-main branch. commit_all accepts these pre-held guards instead of re-acquiring (the queue is non-re-entrant). The merge fork path already holds the queue and self-heals through the shared wrapper. Cross-process in-flight forks remain the documented one-winner-CAS gap. Mechanical prep folded in: mutation IR lowering is hoisted so the touched-table set is known before execution; commit_all gains the held_guards parameter. Flips recreate_over_orphaned_fork_before_cleanup_is_actionable to assert self-heal; fork_collision_with_live_concurrent_fork_is_retryable still holds. Docs: writes.md cancelled-future note, invariants.md cross-process known gap.
reconcile_orphaned_branches keyed orphans on the branch NAME (absent from the manifest), so it only reclaimed forks from a fully-deleted branch. A fork left on a still-live branch by an interrupted first-write was never reclaimed — the backstop the handoff expected cleanup to provide did not cover that case. Broaden it to a per-table authority test: a Lance branch B on table T is an orphan iff B is not a live manifest branch (delete-leftover) OR the manifest's branch-B snapshot does not place T on B (interrupted first-write). Per-branch snapshots are resolved once and cached across tables. Legitimately-forked tables, main, and internal/system branches are never reclaimed; children are dropped before parents to avoid Lance's referenced-parent RefConflict. The commit-graph half stays whole-branch (per-table doesn't apply there). This is the guaranteed-convergence backstop to the write-path self-heal: it reclaims any fork the write path never revisits, and is what Lance's own create_branch docstring asks embedders to provide for zombie/orphan refs.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 37f5af587a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The fork reclaim force-deletes a Lance branch ref, gated on the caller's proof that the manifest does not place the table on the branch. But the first-write path obtains that proof via snapshot_for_branch, which returns the coordinator's CACHED snapshot when the handle is bound to the branch (an embedded handle on the branch, or branch_merge's target swap). If that snapshot is stale and a concurrent writer already published a legitimate fork, the reclaim would force-delete it and re-fork from source, stranding the manifest at a version the recreated ref no longer has. Make the destructive primitive own its safety precondition: re-derive it from a FRESH manifest read (fresh_snapshot_for_branch, which bypasses the cache) immediately before force-deleting. If fresh authority shows the table is on the branch, refuse with a retryable conflict instead of destroying a valid fork. Correct for any caller regardless of snapshot staleness. Also stop branching on Lance's exact RefConflict prose (loosened match; typed-variant is the durable follow-up). Addresses PR review (Codex P1, Greptile P2).
A node delete cascades to every edge table touching that node (execute_delete_ node), forking those edge tables during execution. But touched_table_keys derived the up-front fork-queue set from the IR ops alone (just node:Type), so a branch delete that forks node + cascade edges held only the node queue — commit_all then saw cascade-edge keys it had no guard for. The touched set is a pure function of (IR ops + catalog), so compute the COMPLETE set: op types plus, for delete-node ops, the cascade edges derived the same way the executor derives them (from_type/to_type match). Pre-computed now equals actual by construction. Also promote commit_all's held-guard coverage check out of debug_assert into an all-builds check that fails the write with a typed manifest_internal error: a load-bearing serialization invariant must fail loudly+safely in release, not silently proceed unguarded if a future execution path ever touches a table outside the pre-computed set. Adds branch_cascade_delete_forks_node_and_edges_under_held_queues, which drives the cascade path on a branch (the gap the existing insert/load tests missed). Addresses PR review (Cursor medium, Greptile P2).
The broadened per-table reconciler force_delete'd an orphan candidate on a LIVE branch without holding the per-(table, branch) write queue. An in-process first-write fork in its fork->publish window holds that queue and has not yet advanced the manifest, so it looks exactly like an origin-2 orphan — concurrent cleanup could delete the ref the writer still holds and is about to publish. (The old branch-name-based reconciler did not have this race: a deleted branch cannot have a live first-write.) Bring the reconciler under the same invariant the write-path reclaim already obeys: never force_delete a fork ref without holding the (table, branch) write queue AND confirming, under it, from a fresh read, that the ref is still manifest-unreferenced. Acquire one key at a time (no lock-order inversion vs multi-table acquire_many writers); if the writer published meanwhile, the fresh re-check sees the table on the branch and skips. Cross-process writers remain the documented one-winner-CAS gap. Addresses PR review (Cursor high).
fork_branch_from_state mapped ANY create_branch failure to RefAlreadyExists, routing transient I/O / version / Lance-internal errors into the destructive reclaim path and masking the real error as a retryable conflict. Branch on the actual fact instead: on create_branch failure, check whether the ref exists (list_branches). Only a genuinely pre-existing ref — a fully-formed manifest-unreferenced fork — is a reclaim candidate; any other failure propagates with fidelity. We deliberately do NOT force-delete on a not-found-ref failure: it is indistinguishable from a transient error on a fresh create, and force-deleting there is the overreach the fresh-authority guard already removed. A phase-1-only Lance zombie (rarer; create_branch interrupted mid its two internal phases) surfaces as the propagated error for manual reclaim. Addresses PR review (Cursor medium).
…ive branch The reconcile pre-delete re-check treated ANY fresh_snapshot error as 'still an orphan' and proceeded to force_delete. A transient manifest read failure on a LIVE branch could therefore destroy a fork the manifest still considers legitimate — inconsistent with the write-path reclaim (aborts on the same error) and the candidate scan (skips on snapshot failure). Distinguish the two origins under the queue: a branch absent from the manifest authority (origin 1) is a confirmed orphan and is deleted without a fresh read (no live writer can hold a deleted branch's queue); a LIVE branch (origin 2) gets the fresh re-check and, on a transient read error, is SKIPPED — never destroyed on ambiguity — converging on a later cleanup. Same don't-destroy-on- ambiguous-error principle as the create_branch failure classification. Addresses PR review (Cursor medium).
Consolidates the reconcile/reclaim hardening from PR review (the earlier per-site
commits were collapsed when reconciling with the main merge). Both destructive
fork-ref sites — the write-path reclaim and the cleanup reconciler — now share
one classifier, classify_fork_ref -> ForkRefStatus { Legitimate, Orphan,
Indeterminate }, evaluated from FRESH manifest authority under the held
(table, branch) write queue. A fork ref is destroyed ONLY on a confirmed Orphan;
a Legitimate (concurrent writer published a real fork) or Indeterminate
(transient read) status is never destroyed — the write path maps it to a
retryable conflict, cleanup maps it to skip. This closes, by construction:
- reclaim trusting a possibly-cached caller proof (Codex P1);
- reconcile racing an in-process live fork without the queue (Cursor);
- delete-on-transient-error in the re-check (Cursor/Greptile);
- origin-1 trusting a stale live_branches capture for a created-since branch
(Cursor/Greptile P1).
Having one classifier removes the duplication that let the two sites drift.
ForkOutcome is made pub to match the sealed trait method returning it. Verified
green on Lance 7.0.0 (full engine suite + 48/48 failpoints).
…ghost) Both fork-ref reclaim sites (write-path reclaim + cleanup reconciler) route their destroy/skip decision through classify_fork_ref, but it had no direct test — reverting the fresh-authority logic was not test-detectable. Add a deterministic in-source unit test that forges each state and asserts the status: a manifest-placed fork -> Legitimate (never destroyed); a ref the manifest does not place on the branch -> Orphan; a ref for a branch absent from the manifest -> Orphan (ghost reclaim preserved). This makes the core fresh-authority decision behind every reclaim fix revert-detectable in one place. (The Indeterminate arm — transient read on a live branch -> skip — needs an injected read failure and is left to the failpoints suite; the cross-process cleanup-vs-writer and cached-snapshot reclaim races are the documented one-winner-CAS gap, not reachable same-process bugs, so they are not faked here.)
Closes the last untested classify_fork_ref arm. Adds a 'classify.fresh_read' failpoint (no-op without the failpoints feature) that simulates a transient failure of the fresh-authority read, and a failpoints test driving it through cleanup: a genuine origin-2 orphan on a LIVE branch whose fresh re-check fails classifies as Indeterminate, so the reconciler SKIPS it (never destroys on an inconclusive read) and reclaims it on the next run once the read succeeds. This makes the don't-destroy-on-ambiguity rule revert-detectable end-to-end. The only paths now left untested are the cross-process cleanup-vs-writer and reclaim-vs-publish races — the documented one-winner-CAS gap (cleanup is &mut self / CLI-only, so no reachable same-process race), not faked here.
# Conflicts: # crates/omnigraph/tests/maintenance.rs
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit d80cb12. Configure here.
# Conflicts: # crates/omnigraph-cli/tests/cli_schema_config.rs
ragnorc
added a commit
that referenced
this pull request
Jun 15, 2026
Resolved two conflicts against main (#253 schema-apply actor threading / RFC-011 D10; #231 self-heal manifest-unreferenced branch forks): - schema_apply.rs: my queue-before-snapshot reorder MOVED the touched-table queue-acquisition block to before the snapshot; #231's edit here was a comment-only update to that (now-relocated) block. Kept my relocation; #231's actual reconciler logic lives in recovery.rs/optimize.rs (auto-merged). - invariants.md: add/add in Known Gaps — kept BOTH my sentence narrowing the recovery-serialization gap (the apply publish is now CAS-fenced) AND #231's new "Fork reclaim is in-process-safe only" bullet. Verified the API surfaces this branch calls are unchanged post-merge: write_queue::acquire_many, manifest::commit_changes_with_expected, apply_schema_with_lock (no actor param — #253 enforces in the outer apply_schema, publish stays system-attributed).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

What & why
A first write to a table on a branch lazily forks it via Lance
create_branch— a durable, two-phase op that advances Lance state before the atomic manifest publish — so a crash, deploy restart, or cancelled request future between the fork and the publish left a fully-formed branch ref the manifest never referenced, wedging every subsequent write to that table on that branch with"orphaned table state … incomplete prior delete; run cleanup"(andcleanupcouldn't even fix it, since the branch was still live). This makes the per-table fork derived state of the manifest: the fork now returns a typedForkOutcomeand self-heals a manifest-unreferenced leftover (reclaim + re-fork under the write queue) on the next write across load/mutate/merge, withcleanup's reconciler broadened to a per-table authority test as the guaranteed backstop.Backing issue / RFC
Checklist
chore:commitwrites.rs), reconcile coverage (maintenance.rs), failpoints flip + preserved retryable testwrites.md,invariants.md) updatedNotes for reviewers
Reclaim is safe because branch first-writes now acquire the per-
(table, branch)write queues up front (held through publish), so no in-process writer can be mid-fork; cross-process in-flight forks remain the already-documented one-winner-CAS gap (noted ininvariants.md). Commits are file-granular and ordered red→green (3edbe2eis the RED test,6e3204dthe fix) since interactive hunk staging wasn't available. Only steady-state cost is added serialization of concurrent first-writes to the same(table, branch)— which already conflict at commit; main-branch and already-forked writes are unaffected.Note
High Risk
Changes destructive fork reclaim, manifest authority checks, and write-queue ordering on load/mutate/merge/cleanup paths; mistakes could delete live forks or leave writes unserialized, though tests and failpoints target those cases.
Overview
Fixes a durability bug where an interrupted first-write fork (Lance
create_branchbefore manifest publish) left a branch ref the manifest never referenced, wedging later writes with “run cleanup” even on live branches.Write path:
fork_branch_from_statenow returnsForkOutcome::RefAlreadyExistsinstead of a hard error; the engine runsreclaim_orphaned_fork_and_reforkafterclassify_fork_refconfirms the ref is an orphan from fresh manifest authority. Load and mutate acquire per-(table, branch)write queues up front for all touched tables (including cascade-delete edges) and pass held guards intocommit_allso reclaim cannot race in-process forks and cannot self-deadlock on re-acquire.Cleanup:
reconcile_orphaned_branchestreats orphans per table (live branch but table not placed on it), re-validates under the write queue via the same classifier, and skips onIndeterminatereads.Dev docs and tests are updated (self-heal regressions, failpoint flips); minor comment/README tidy.
Reviewed by Cursor Bugbot for commit 1eea371. Bugbot is set up for automated code reviews on this repo. Configure here.
Greptile Summary
This PR fixes a durability wedge on the branch first-write path: Lance's
create_branchadvances storage state before the manifest publish, so a crash or cancelled future left a fully-formed branch ref the manifest never referenced, causing every subsequent write to that table on that branch to error with "incomplete prior delete; run cleanup." The fix makes the fork a self-healing derived-state operation:fork_branch_from_statenow returns a typedForkOutcomerather than surfacing a collision as an error,reclaim_orphaned_fork_and_reforkforce-deletes and re-forks under fresh manifest authority, andreconcile_orphaned_branchesis broadened to a per-table authority test (origin-2 orphans on live branches).mutation.rsandloader/mod.rspre-acquire per-(table, branch)write queues for all touched tables (including cascade edges for node deletes) before the first fork and hold them through the publish;commit_allpromotes the guard-coverage check fromdebug_assertto an all-builds error.reconcile_orphaned_branchesnow tests each(table, branch)pair against fresh manifest authority under the write queue via the sharedclassify_fork_refclassifier (Orphan/Legitimate/Indeterminate), preventing the stale-live_branchesshortcut from destroying a legitimately-published fork on a newly-created branch.writes.rs, per-table cleanup coverage inmaintenance.rs, and failpoint tests pinning theIndeterminateskip, reconcile converge, and snapshot-failure caching paths infailpoints.rs.Confidence Score: 5/5
Safe to merge. The destructive fork reclaim requires two independent gates to agree (in-process write queue serialization + fresh manifest authority via classify_fork_ref returning Orphan) before any force-delete runs; the Indeterminate path consistently skips rather than destroys on ambiguity. All previously flagged issues have been addressed.
The core fix is correct by construction: ForkOutcome replaces an opaque error with a typed signal, classify_fork_ref consolidates the 'safe to destroy?' decision into one site shared by both reclaim paths, and the write-queue pre-acquisition closes the in-process race. Previous review findings (debug_assert promotion, string-based error matching, Err(_) => destructive action, stale live_branches shortcut) are all addressed. The remaining inline comment concerns a rare diagnostic edge case in the Legitimate error arm where expected and actual versions could coincide — retryable and safe, just slightly misleading. Test coverage spans red→green regression, failpoint-injected Indeterminate/skip/converge paths, and the new per-table cleanup authority test.
table_ops.rs (reclaim_orphaned_fork_and_refork) — the Legitimate error arm's second fresh_snapshot_for_branch call can produce an expected==actual version mismatch if the read fails between classify and the version lookup; suggested fix inline.
Important Files Changed
Sequence Diagram
sequenceDiagram participant Caller as mutate_as / load_as participant WQ as WriteQueue participant OB as open_owned_dataset_for_branch_write participant FS as fork_dataset_from_entry_state (Omnigraph wrapper) participant TS as TableStore.fork_branch_from_state participant CFR as classify_fork_ref participant ROR as reclaim_orphaned_fork_and_refork participant CA as commit_all Caller->>WQ: acquire_many(touched tables x branch) [if needs_fork] WQ-->>Caller: guards held Caller->>OB: open for branch write OB->>FS: fork (source to active_branch) FS->>TS: create_branch(target_branch, source_version) alt RefAlreadyExists TS-->>FS: ForkOutcome::RefAlreadyExists FS->>ROR: reclaim_orphaned_fork_and_refork ROR->>CFR: classify_fork_ref(fresh authority) CFR-->>ROR: Orphan / Legitimate / Indeterminate alt Orphan ROR->>TS: force_delete_branch + re-fork ROR-->>FS: Ok(SnapshotHandle) else Legitimate / Indeterminate ROR-->>Caller: retryable error end else Created TS-->>FS: ForkOutcome::Created(dataset) FS-->>OB: Ok(SnapshotHandle) end OB-->>Caller: (dataset, Some(active_branch)) Caller->>CA: commit_all(held_guards) CA->>CA: coverage check (queue_keys subset acquired_keys) CA-->>Caller: (updates, versions) Note over Caller,WQ: guards released at end of scopeComments Outside Diff (1)
crates/omnigraph/src/exec/mutation.rs, line 456-469 (link)needs_fork = falseleaves a fork-without-queue gap on cross-process branch recreateIf
needs_forkisfalseat snapshot time (all touched tables appear already-forked), no guards are pre-acquired. If a foreign-process then deletes and recreates the branch between this snapshot fetch andexecute_named_mutation, some tables become unforked;open_owned_dataset_for_branch_writere-detects this and callsdb.fork_dataset_from_entry_state, which may reachreclaim_orphaned_fork_and_refork— without the per-(table, branch)queue held. The invariant comment onreclaim_orphaned_fork_and_reforkstates callers "MUST already hold" the queue.This is acknowledged in
invariants.mdas the one-winner-CAS gap and leads to retries rather than corruption, so it's not a blocking concern. Worth calling out explicitly since the comment onreclaim_orphaned_fork_and_reforkreads as an unconditional precondition but actually has this cross-process carve-out.Reviews (13): Last reviewed commit: "Merge remote-tracking branch 'origin/mai..." | Re-trigger Greptile