Repository navigation
Conversation
See context/import-peer-id-reuse.md for comparison and allocation rationale. Co-Authored-By: GPT-6 (Codex) <noreply@openai.com>
Contributor
WASM Size Report
|
Co-Authored-By: GPT-6 (Codex) <noreply@openai.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reduce repeated-history import costs without narrowing peer-id reuse detection (#1118):
Rationale
The previous overlap check parsed and retained every cold local block. Snapshot decoding also allocated incoming known text/list values in the document arena and cloned changes that trimming dropped. This change keeps the full comparison scope: it skips only byte-proven equal blocks and otherwise checks the same op content/dependencies. Existing limits (delta sync without the conflicting prefix, no-new-op duplicates, shallow history) are unchanged and documented in
context/import-peer-id-reuse.md.The cold comparison runs only when an imported change reaches the end of the local block. Short imported changes over a non-identical cold block still parse it, so the remaining roughly 2.2x
overlap_snapcost atFAT=1is expected. The cold reader is an optimization, not a second validity rule: only the ordinary parser may declare local history unparsable.Under
#[cfg(test)], the override at the end ofpreflight_import_changesforces rollback only forapplies_to_dag && !detached; detached tests continue to exercise the production rollback decision.Measurements
Measured on macOS arm64,
rustc 1.96.0 (ac68faa20 2026-05-25), release builds,CARGO_BUILD_JOBS=4. Baselines: 1.16.3ad5b2a6d4546473d9f4a96412d2ae6808c8403ec,main
c00c9fa501f8d32f68d6255eacb7035a67fb6ab6, fixed = optimization commit0c8253e1b966d475a3fb5376a9cdb769ea0d0b40. The table retains the initialPR measurements; the review follow-up reran correctness checks, not this matrix.
The verifier directory was read only; its probe and NOSEED/SEEDSYNC patch were
copied here. Each cell has 3 interleaved old/main/fixed process runs, each reporting
3 timed rounds after a warmup. Tables show medians of process medians.
The machine is shared; recorded one-minute load ranged 3.48-4.37.
Use ratios, not an individual wall time.
The first five scenarios use
FAT=64 REPEATS=4. Other scenarios useFAT=1;text_streamusesNOSEED=1(no concurrent seed).snap_memhas no warmup:its time column is the median of each process's four growing-snapshot imports,
then the median over 3 processes. Its first-import latency is separately covered
by
snap_plus.Retained allocator memory
Process-wide
malloc_zone_statistics().size_in_use, in KiB.snap_plusis theone-import delta;
snap_memis the delta from the loaded base to import 4.These statistics have allocator growth steps and include state/history caches;
the exact arena regression test verifies that no incoming known prefix survives.
Fixed
snap_memretained-heap deltas from its loaded base after imports 1/2/3/4(median of 3 processes, KiB):
FAT=1 controls for the original regression shape
Same release/interleaving/round/repeat settings; these reproduce the originally
reported 9 ms and 15 ms regressions without large inserted strings.
R1 retains some comparison cost above 1.16.3, which did not check known history.
Movable-list element validation is unchanged:
mlist_batchretains its roughly1.2x cost relative to 1.16.3. Text streaming and batch controls are comparable to
main. Snapshot decoding is faster than both baselines and repeated imports plateau.
Reproduce with separate binaries named
old,main,fixed:Samples: FAT=64 matrix,
FAT=1 controls.
The runner also saves full probe output and load readings to
raw.jsonl.Validation
Re-run after the review fixes, with
CARGO_BUILD_JOBS=4:cargo test -p loro --test import_reused_peer_id: 11 passed.cargo test -p loro-internal --lib cold_text: 3 passed (the new positive/fallback tests and the cold dependency-conflict test).cargo test -p loro-internal --lib import_atomicity: 18 passed.cargo test -p loro-internal --lib known_history: 2 passed.cargo test -p loro-internal --lib oplog::change_store: 23 passed, including this PR's identical-block, snapshot-arena, range-overflow and cold-block tests.cargo test -p loro-internal --lib: 415 passed, 4 ignored.pnpm test: 1754 passed, 37 skipped across 102 nextest binaries; 62 doctests passed (57 + 1 + 4).pnpm check: passed (cargo clippy --all-features -- -Dwarnings).pnpm test-loom: 9 passed, 0 failed (LOOM_MAX_PREEMPTIONS=2, release,RUSTFLAGS='--cfg loom').git diff --check: passed.The initial release probe
checkalso passed: 15 PASS + CHECK_OK on old/main/fixed; old still accepts the deliberate peer-id conflict, main/fixed reject it as expected.No API or binary-format change. No broad fuzz/browser matrix was run. The verifier worktree was not modified.