Skip to content

3/5: Relocate garbage-heavy blob files without SST rewrites (#15307) - #15307

Open
xingbowang wants to merge 3 commits into
facebook:mainfrom
xingbowang:export-D123122449
Open

xingbowang wants to merge 3 commits into
facebook:mainfrom
xingbowang:export-D123122449

Conversation

@xingbowang

@xingbowang xingbowang commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary:
GitHub Issue: #15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449

@meta-cla meta-cla Bot added the CLA Signed label Oct 2, 2026
@meta-codesync

meta-codesync Bot commented Oct 2, 2026

Copy link
Copy Markdown

@xingbowang has exported this pull request. If you are a Meta employee, you can view the originating Diff in D123122449.

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

⚠️ clang-tidy: 5 warning(s) on changed lines

Completed in 1503.3s.

Summary by check

Check Count
bugprone-argument-comment 5
Total 5

Details

db/blob/blob_file_addition_test.cc (4 warning(s))
db/blob/blob_file_addition_test.cc:122:29: warning: argument name 'origin_file_number' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:123:29: warning: argument name 'carrier_file_size' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:148:29: warning: argument name 'origin_file_number' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/blob/blob_file_addition_test.cc:149:29: warning: argument name 'carrier_file_size' in comment does not match parameter name 'value' [bugprone-argument-comment]
db/db_impl/db_impl_compaction_flush.cc (1 warning(s))
db/db_impl/db_impl_compaction_flush.cc:460:57: warning: argument name 'sequence' in comment does not match parameter name 's' [bugprone-argument-comment]

@meta-codesync meta-codesync Bot changed the title 3/4: Relocate garbage-heavy blob files without SST rewrites 3/4: Relocate garbage-heavy blob files without SST rewrites (#15307) Oct 3, 2026
xingbowang added a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
xingbowang pushed a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
xingbowang pushed a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
xingbowang added a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
xingbowang pushed a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
xingbowang added a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
Xingbo Wang added 2 commits October 3, 2026 08:54
Summary:
Pull Request resolved: facebook#15306
GitHub Issue: facebook#15301

Why the change is required:
Relocating a blob without rewriting every referring SST requires a logical identity that remains stable while the physical carrier changes.

Previous behavior:
BlobIndex encoded only a physical blob file number and offset. Blob metadata and VersionStorageInfo therefore assumed every SST reference named the current physical file.

New behavior:
The default-off enable_blob_indirection option makes new blob files logical identity roots. Their BlobIndexes encode the stable origin file number, origin offset, value size, checksum, and compression. BlobFileAddition and BlobFileMetaData persist whether a physical file is the identity or replacement carrier for that origin, and VersionStorageInfo maintains the current origin-to-physical route across MANIFEST replay and snapshots.

Implementation:
The change adds the indirect BlobIndex encoding and checksum, immutable option plumbing and validation, generated C API accessors and round-trip coverage, identity metadata in blob additions, origin-indexed Version construction and snapshot serialization, plus post-recovery rejection of unsupported configurations. BlobFileBuilder emits stable references only when the option is enabled.

Safety, compatibility, and limitations:
The option is false by default, so existing direct BlobIndexes and ordinary blob files are unchanged. The new MANIFEST metadata is intentionally forward-incompatible so older binaries fail closed. Blob direct write, blob compression, TTL integration, and remote compaction remain unsupported. This diff establishes identity and routing; carrier reads and relocation are added by later diffs in the stack.

Validation:
The stack boundary compiles independently. Final-stack coverage exercises index encoding, option and C API persistence, MANIFEST replay, Version construction, carrier reads and relocation, and existing direct-reference behavior. The normal and checked-Status full test suites pass.

Differential Revision: D123122447
Summary:
Pull Request resolved: facebook#15308
GitHub Issue: facebook#15301

Why the change is required:
A stable BlobIndex needs an indexed physical carrier that can translate an origin offset without loading a custom map or rewriting referring SSTs.

Previous behavior:
Blob reads opened blob-log files and interpreted offsets directly. RocksDB had embedded-blob block-based tables, but no carrier discriminator, persisted logical origin, or BlobSource path for origin-offset lookup.

New behavior:
A block-based table can act as a Blob GC carrier keyed by the fixed-width original blob offset. Each entry points to a colocated embedded value. Scalar, MultiGet, whole-value lazy, cache-only, and asynchronous reads resolve the stable origin through the Version route and perform an on-demand table lookup. Whole-value reads validate the stable checksum when requested. Strict lazy sub-range reads preserve the existing no-read-amplification policy and do not read or verify the whole value unless force_verify escalates them to a whole read.

Implementation:
TableBuilderOptions carries the logical origin, BlockBasedTableBuilder writes a carrier prefix and origin property, and shared helpers encode carrier keys and derive restricted table options. BlobFileReader opens carriers through the existing table reader and block cache, validates their size, properties, origin, protection settings, lookup key, and embedded index, and reuses the same-file blob path for the payload. AdaptiveTableFactory exposes nested block-based reader options.

Safety, compatibility, and limitations:
Ordinary table builds and direct BlobIndexes retain their existing paths. Carrier lookup requires block-based format version 7 or newer with bytewise ordering and binary-search indexes. Missing, stale, foreign, truncated, or structurally corrupt carrier metadata fails closed. Payload verification follows the read-mode semantics above; strict sub-range reads do not promise whole-value stable-checksum verification without force_verify. This diff reads carriers but does not yet create or publish replacements.

Validation:
The stack boundary compiles independently. Final-stack coverage exercises cold and cached carrier opens, scalar and batched reads, range reads, corruption checks, cache-only behavior, legacy direct reads, and the normal and checked-Status full test suites.

Differential Revision: D123122448
xingbowang pushed a commit to xingbowang/rocksdb that referenced this pull request Oct 3, 2026
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
…#15307)

Summary:
Pull Request resolved: facebook#15307
GitHub Issue: facebook#15301

Why the change is required:
FIFO blob GC couples reclamation to blob age and SST compaction. That rewrites SSTs even when stable blob identities make a physical-only replacement possible.

Previous behavior:
Live blobs were relocated only while compacting referring SSTs, and the standalone prototype depended on a custom blob-map container.

New behavior:
When blob indirection and blob garbage collection are enabled, RocksDB selects an eligible garbage-heavy origin, proves its complete live reference set in a pinned visible-key scan, writes a smaller embedded-table carrier, syncs and verifies it, and atomically publishes the new origin route in the MANIFEST without modifying SSTs. Repeated GC can replace a carrier again; eligibility never compares logical bytes containing original user keys with the carrier's different physical representation.

Implementation:
The change adds garbage-ratio candidate scoring, single-origin census collection, embedded carrier construction, source revalidation, serialized MANIFEST publication, route replacement and rollback, scheduling and retry handling, disk-space reservation, listener accounting, recovery validation, incomplete-output quarantine, obsolete-file cleanup, and repeated relocation support. The table builder's exact output size is the authority for whether a replacement reclaims physical bytes.

Safety, compatibility, and limitations:
The path remains default-off and leaves ordinary compaction priority unchanged. It rejects incomplete censuses, compressed blobs, merge-operator bases that cannot be exposed, non-block-based tables, old table formats, TTL files, blob direct write, and remote compaction. A carrier is published only after Finish, Sync, Close, rename, directory sync, source revalidation, and a physical-size check prove it is durable and smaller. Pinned Versions retain superseded carriers until normal obsolete-file processing can remove them.

Validation:
This single-origin boundary compiles independently. Final-stack coverage exercises relocation without SST rewrites, repeated relocation with both small and 8-KiB user keys, recovery and replay, cancellation, concurrent route changes, column-family drop, corruption, resource bounds, obsolete-file handling, and both full test modes.

Differential Revision: D123122449
@meta-codesync meta-codesync Bot changed the title 3/4: Relocate garbage-heavy blob files without SST rewrites (#15307) 3/5: Relocate garbage-heavy blob files without SST rewrites (#15307) Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant