The parquet writer in milvus-storage never calls dictionary_pagesize_limit,
so it inherits parquet's 1 MB default. Combined with our byte-bounded row
group size (DEFAULT_MAX_ROW_GROUP_SIZE = 1 MB), this means a single column's
dictionary can grow large enough to dominate the row group before the writer
falls back to PLAIN.
For high-cardinality columns, the dictionary page becomes a redundant copy
of the unique values, since each row group starts a fresh dictionary.
Measured impact on a 300k-row benchmark with high-cardinality random data
(parquet, zstd):
file size: 73.6 MB -> 46.5 MB (-37%)
write time: 6.0 s -> 0.9 s (6.6x faster)
read time: 210 ms -> 181 ms (+16% throughput)
Per-column decomposition at the default 1 MB limit shows the dictionary
page accounts for 60-80% of each column chunk's bytes on this workload.
Proposed fix: hardcode dictionary_pagesize_limit to row_group_size / 16
(64 KB) in convert_write_properties. This triggers fallback early on
high-cardinality columns while preserving dictionary encoding for columns
whose dictionary never grows past 64 KB.
The parquet writer in milvus-storage never calls dictionary_pagesize_limit,
so it inherits parquet's 1 MB default. Combined with our byte-bounded row
group size (DEFAULT_MAX_ROW_GROUP_SIZE = 1 MB), this means a single column's
dictionary can grow large enough to dominate the row group before the writer
falls back to PLAIN.
For high-cardinality columns, the dictionary page becomes a redundant copy
of the unique values, since each row group starts a fresh dictionary.
Measured impact on a 300k-row benchmark with high-cardinality random data
(parquet, zstd):
file size: 73.6 MB -> 46.5 MB (-37%)
write time: 6.0 s -> 0.9 s (6.6x faster)
read time: 210 ms -> 181 ms (+16% throughput)
Per-column decomposition at the default 1 MB limit shows the dictionary
page accounts for 60-80% of each column chunk's bytes on this workload.
Proposed fix: hardcode dictionary_pagesize_limit to row_group_size / 16
(64 KB) in convert_write_properties. This triggers fallback early on
high-cardinality columns while preserving dictionary encoding for columns
whose dictionary never grows past 64 KB.