Commit 49c0313
committed
perf: SIMD short-circuit in JoinHashMap probe
[AURON-2160] Optimize join hash map probe by checking hash_matched
first before computing empty mask. This reduces ~50% SIMD instructions
when hash hit rate is high (typical join scenarios).
Before: Always compute both hash_matched and empty SIMD masks.
After: Only compute empty mask when hash_matched has no hits.
Also add a criterion microbenchmark (benches/join_hash_map.rs) covering
realistic BHJ build sizes (5M/10M/20M keys) × three hit rates (0/50/100%).
Results on Apple M2 Pro (probe_size=4096):
build size | hit=0% | hit=50% | hit=100%
----------------+---------+---------+---------
5M (~128 MB) | 6.63 µs | 6.52 µs | 6.35 µs
10M (~256 MB) | 6.68 µs | 6.50 µs | 6.36 µs
20M (~512 MB) | 6.70 µs | 6.59 µs | 6.36 µs
Latency stays flat because prefetch_read_data (4-step ahead) fully
pipelines cache misses. The hit=100% path is consistently ~4-5% faster,
aligning with the optimization goal. Instruction-count savings can be
confirmed on x86 via: perf stat -e instructions
Run benchmark:
cargo bench --bench join_hash_map -p datafusion-ext-plans1 parent 016c7a4 commit 49c0313
5 files changed
Lines changed: 357 additions & 24 deletions
File tree
- native-engine/datafusion-ext-plans
- benches
- src/joins
0 commit comments