Skip to content

[Enhancement]: parquet dict_pagesize_limit not tuned for 1 MB row groups #543

Description

@tedxu

The parquet writer in milvus-storage never calls dictionary_pagesize_limit,
so it inherits parquet's 1 MB default. Combined with our byte-bounded row
group size (DEFAULT_MAX_ROW_GROUP_SIZE = 1 MB), this means a single column's
dictionary can grow large enough to dominate the row group before the writer
falls back to PLAIN.

For high-cardinality columns, the dictionary page becomes a redundant copy
of the unique values, since each row group starts a fresh dictionary.

Measured impact on a 300k-row benchmark with high-cardinality random data
(parquet, zstd):

file size: 73.6 MB -> 46.5 MB (-37%)
write time: 6.0 s -> 0.9 s (6.6x faster)
read time: 210 ms -> 181 ms (+16% throughput)

Per-column decomposition at the default 1 MB limit shows the dictionary
page accounts for 60-80% of each column chunk's bytes on this workload.

Proposed fix: hardcode dictionary_pagesize_limit to row_group_size / 16
(64 KB) in convert_write_properties. This triggers fallback early on
high-cardinality columns while preserving dictionary encoding for columns
whose dictionary never grows past 64 KB.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions