3/5: Relocate garbage-heavy blob files without SST rewrites (#15307) - #15307
Open
xingbowang wants to merge 3 commits into
Open
xingbowang wants to merge 3 commits into
xingbowang wants to merge 3 commits into
Conversation
|
@xingbowang has exported this pull request. If you are a Meta employee, you can view the originating Diff in D123122449. |
|
| Check | Count |
|---|---|
bugprone-argument-comment |
5 |
| Total | 5 |
Details
db/blob/blob_file_addition_test.cc (4 warning(s))
db/blob/blob_file_addition_test.cc:122:29: warning: argument name 'origin_file_number' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:123:29: warning: argument name 'carrier_file_size' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:148:29: warning: argument name 'origin_file_number' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:149:29: warning: argument name 'carrier_file_size' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/db_impl/db_impl_compaction_flush.cc (1 warning(s))
db/db_impl/db_impl_compaction_flush.cc:460:57: warning: argument name 'sequence' in comment does not match parameter name 's' [bugprone-argument-comment]
xingbowang
force-pushed
the
export-D123122449
branch
from
October 3, 2026 11:47
5b5ed75 to
8e418c5
Compare
xingbowang
added a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
pushed a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
pushed a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
added a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
force-pushed
the
export-D123122449
branch
from
October 3, 2026 12:20
8e418c5 to
baf67eb
Compare
xingbowang
pushed a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
added a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
force-pushed
the
export-D123122449
branch
from
October 3, 2026 13:11
baf67eb to
91e52f7
Compare
added 2 commits
October 3, 2026 08:54
Summary: Pull Request resolved: facebook#15306 GitHub Issue: facebook#15301 Why the change is required: Relocating a blob without rewriting every referring SST requires a logical identity that remains stable while the physical carrier changes. Previous behavior: BlobIndex encoded only a physical blob file number and offset. Blob metadata and VersionStorageInfo therefore assumed every SST reference named the current physical file. New behavior: The default-off enable_blob_indirection option makes new blob files logical identity roots. Their BlobIndexes encode the stable origin file number, origin offset, value size, checksum, and compression. BlobFileAddition and BlobFileMetaData persist whether a physical file is the identity or replacement carrier for that origin, and VersionStorageInfo maintains the current origin-to-physical route across MANIFEST replay and snapshots. Implementation: The change adds the indirect BlobIndex encoding and checksum, immutable option plumbing and validation, generated C API accessors and round-trip coverage, identity metadata in blob additions, origin-indexed Version construction and snapshot serialization, plus post-recovery rejection of unsupported configurations. BlobFileBuilder emits stable references only when the option is enabled. Safety, compatibility, and limitations: The option is false by default, so existing direct BlobIndexes and ordinary blob files are unchanged. The new MANIFEST metadata is intentionally forward-incompatible so older binaries fail closed. Blob direct write, blob compression, TTL integration, and remote compaction remain unsupported. This diff establishes identity and routing; carrier reads and relocation are added by later diffs in the stack. Validation: The stack boundary compiles independently. Final-stack coverage exercises index encoding, option and C API persistence, MANIFEST replay, Version construction, carrier reads and relocation, and existing direct-reference behavior. The normal and checked-Status full test suites pass. Differential Revision: D123122447
Summary: Pull Request resolved: facebook#15308 GitHub Issue: facebook#15301 Why the change is required: A stable BlobIndex needs an indexed physical carrier that can translate an origin offset without loading a custom map or rewriting referring SSTs. Previous behavior: Blob reads opened blob-log files and interpreted offsets directly. RocksDB had embedded-blob block-based tables, but no carrier discriminator, persisted logical origin, or BlobSource path for origin-offset lookup. New behavior: A block-based table can act as a Blob GC carrier keyed by the fixed-width original blob offset. Each entry points to a colocated embedded value. Scalar, MultiGet, whole-value lazy, cache-only, and asynchronous reads resolve the stable origin through the Version route and perform an on-demand table lookup. Whole-value reads validate the stable checksum when requested. Strict lazy sub-range reads preserve the existing no-read-amplification policy and do not read or verify the whole value unless force_verify escalates them to a whole read. Implementation: TableBuilderOptions carries the logical origin, BlockBasedTableBuilder writes a carrier prefix and origin property, and shared helpers encode carrier keys and derive restricted table options. BlobFileReader opens carriers through the existing table reader and block cache, validates their size, properties, origin, protection settings, lookup key, and embedded index, and reuses the same-file blob path for the payload. AdaptiveTableFactory exposes nested block-based reader options. Safety, compatibility, and limitations: Ordinary table builds and direct BlobIndexes retain their existing paths. Carrier lookup requires block-based format version 7 or newer with bytewise ordering and binary-search indexes. Missing, stale, foreign, truncated, or structurally corrupt carrier metadata fails closed. Payload verification follows the read-mode semantics above; strict sub-range reads do not promise whole-value stable-checksum verification without force_verify. This diff reads carriers but does not yet create or publish replacements. Validation: The stack boundary compiles independently. Final-stack coverage exercises cold and cached carrier opens, scalar and batched reads, range reads, corruption checks, cache-only behavior, legacy direct reads, and the normal and checked-Status full test suites. Differential Revision: D123122448
xingbowang
pushed a commit
to xingbowang/rocksdb
that referenced
this pull request
Oct 3, 2026
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
…#15307) Summary: Pull Request resolved: facebook#15307 GitHub Issue: facebook#15301 Why the change is required: FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible. Previous behavior: Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container. New behavior: When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation. Implementation: The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes. Safety, compatibility, and limitations: The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them. Validation: This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes. Differential Revision: D123122449
xingbowang
force-pushed
the
export-D123122449
branch
from
October 3, 2026 16:08
91e52f7 to
f8f7715
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary:
GitHub Issue: #15301
Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.
Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.
New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.
Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.
Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.
Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.
Differential Revision: D123122449