Repository navigation
PiPNN 2/6: add numerical kernels - #1287
Conversation
There was a problem hiding this comment.
Pull request overview
This PR adds the first set of PiPNN “kernel” building blocks to the DiskANN Rust workspace: SIMD-accelerated top‑k selection for partition assignment and leaf neighbor selection, along with supporting SIMD division and a new lower-triangular A·Aᵀ helper in diskann-linalg.
Changes:
- Add a new
diskann-pipnncrate withpartition_kernelandleaf_kernelimplementations plus extensive correctness tests and Criterion benchmarks. - Extend
diskann-wideto supportDivon relevant f32 SIMD types (native, doubled, and scalar/emulated) and add a corresponding division test macro. - Add
diskann_linalg::sgemm_aat_lower(lower-triangle-only AAT) and wire new crate/tests/CI/mutants exclusions into the workspace.
Reviewed changes
Copilot reviewed 26 out of 27 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| diskann-wide/src/test_utils/ops.rs | Adds test_div! macro to validate lane-wise SIMD division correctness. |
| diskann-wide/src/emulated.rs | Adds Div for scalar/emulated Emulated<f32, N, A> to support division in scalar dispatch. |
| diskann-wide/src/doubled.rs | Adds Div for Doubled<T> to support composite SIMD widths. |
| diskann-wide/src/arch/x86_64/v4/f32x8_.rs | Adds AVX Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v4/f32x4_.rs | Adds SSE Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v4/f32x16_.rs | Adds AVX-512 Div op mapping + division tests. |
| diskann-wide/src/arch/x86_64/v3/f32x8_.rs | Adds AVX Div op mapping + division tests for V3. |
| diskann-wide/src/arch/x86_64/v3/f32x4_.rs | Adds SSE Div op mapping + division tests for V3. |
| diskann-wide/src/arch/x86_64/v3/f32x16_.rs | Adds division tests for the f32x16 V3 path (likely via doubled composition). |
| diskann-wide/src/arch/aarch64/f32x4_.rs | Adds Neon Div op mapping + division tests. |
| diskann-wide/src/arch/aarch64/f32x2_.rs | Adds Neon Div op mapping + division tests. |
| diskann-pipnn/tests/partition_kernel.rs | New integration tests for partition top‑k dispatch correctness and edge cases. |
| diskann-pipnn/tests/leaf_kernel.rs | New integration tests for leaf neighbor top‑k dispatch correctness and edge cases. |
| diskann-pipnn/src/partition_kernel/tests.rs | New unit tests comparing scalar reference vs runtime dispatch and metric contracts. |
| diskann-pipnn/src/partition_kernel.rs | New partition-assignment distance + top‑k kernel with validation and SIMD dispatch. |
| diskann-pipnn/src/lib.rs | New crate root exporting PiPNN kernel modules. |
| diskann-pipnn/src/leaf_kernel/tests.rs | New unit tests for scalar reference parity and workspace behavior. |
| diskann-pipnn/src/leaf_kernel.rs | New fused lower-triangle leaf neighbor kernel with SIMD dispatch and workspace support. |
| diskann-pipnn/Cargo.toml | Defines new diskann-pipnn crate, dev-deps, and benches. |
| diskann-pipnn/benches/kernels.rs | Adds benchmarks for partition top‑k, lower AAT, leaf top‑k, and full leaf workflow. |
| diskann-linalg/tests/sgemm_aat_lower.rs | New tests for lower-triangle AAT behavior and validation errors. |
| diskann-linalg/src/lib.rs | Adds public sgemm_aat_lower API with dimension checks. |
| diskann-linalg/src/faer.rs | Implements sgemm_aat_lower_impl using Faer triangular matmul. |
| Cargo.toml | Adds diskann-pipnn to workspace members and workspace dependencies. |
| Cargo.lock | Records the new diskann-pipnn package entry. |
| .github/workflows/ci.yml | Adds diskann-pipnn to CI test package lists. |
| .cargo/mutants.toml | Adds mutation-test exclusions for kernel code paths and equivalent transformations. |
Comments suppressed due to low confidence (2)
diskann-pipnn/src/leaf_kernel.rs:651
- Same issue as the L2 arm: using
max_simdfor lower clamping can erase NaNs on the Scalar/Emulated backend, making NaN distances rankable. Clamp withlt_simd+selectto preserve NaNs consistently.
Metric::CosineNormalized => {
let distance = F::splat(arch, 1.0) - dot;
zero.max_simd(distance)
}
diskann-pipnn/src/leaf_kernel.rs:664
- The cosine path also uses
zero.max_simd(distance)for clamping, which can collapse NaNs to zero on the Scalar/Emulated backend (viaf32::max). That contradicts the comment about preserving non-rankable NaNs and can change output ordering. Prefer anlt_simd+selectclamp here as well.
let distance = one - cosine;
// Comparisons with NaN are false, so this explicit lower clamp
// preserves non-rankable NaNs while matching the existing PiPNN
// distance formulas for finite values.
zero.max_simd(distance)
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
e204cb9 to
b046174
Compare
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1287 +/- ##
==========================================
+ Coverage 91.55% 91.61% +0.06%
==========================================
Files 521 567 +46
Lines 100302 112317 +12015
==========================================
+ Hits 91828 102904 +11076
- Misses 8474 9413 +939
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (2)
diskann-pipnn/src/partition_kernel/tests.rs:20
- The
PartitionTopKcontract forMetric::L2expectsleader_scalesto contain squared leader norms (see docs anddistance(Metric::L2, ..)test). This helper currently populates unsquared norms, which makes the test data inconsistent with the public API contract and could hide contract-related bugs.
let leader_scales = match metric {
Metric::L2 => (0..leaders).map(|leader| (leader + 1) as f32).collect(),
Metric::Cosine => (0..leaders)
.map(|leader| {
diskann-pipnn/src/partition_kernel.rs:61
InvalidFanout’s error message says the maximum is{maximum}, but validation also rejectsfanout > leaders. Whenleaders < maximumthis message is misleading (it implies the only limit is{maximum}). Consider spelling out both constraints in the message so callers immediately see why it failed.
#[error("invalid fanout {fanout} for {leaders} leaders; maximum is {maximum}")]
8fb4e92 to
20ab8a0
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 25 out of 26 changed files in this pull request and generated no new comments.
Suppressed comments (1)
diskann-pipnn/src/partition_kernel.rs:294
- For
Metric::Cosine, NaN norms currently produce a finite distance (1.0) becausedenominator.gt_simd(0)is false for NaN, so the lane falls back tocosine = 0. That makes NaN-derived pairs/leaders “rankable”, which contradicts the module’s stated NaN-rejection behavior and differs fromdiskann-vectorcosine semantics (NaN norms propagate to a NaN similarity/distance). Consider explicitly preserving NaN denominators so the resulting distance stays NaN and is ignored byinsert_topk.
let denominator = row_norm * leader_norm;
let valid = denominator.gt_simd(zero);
let safe_denominator = valid.select(denominator, one);
let cosine = valid.select(dot / safe_denominator, zero);
one - cosine
Aditya Krishnan (arkrishn94)
left a comment
There was a problem hiding this comment.
Thanks Weiyao, this is progress from the previous mega-PR. I still have some big-picture comments (we covered most of these offline) -
- Documentation: As I mentioned, we need thorough documentation in the
diskann-pipnncrate. The main modules,partition_kernelandleaf_kernelneed documentation up top, highlighting the main structures and how they are used - e.g.process_rows_binary/unaryandnearest_leaders. Similarly withprocess_pairs_simd_*andnearest_leaf_neighbors - Testing: I am concerned about the lack of testing for partition_kernel.rs and
leaf_kernel.rs.- I notice some e2e integration tests but these kernels should be thoroughly tested, sweeping different input parameters, architectures and edge cases. This is especially needed given the amount of unsafe code.
- That brings me to miri - there should be miri tests too.
- I'm curious why are the tests in a separate submodule to the main files (for partition_kernel.rs and leaf_kernel.rs)? Let's try to keep tests along with the code being tested.
- Criterion: Since criterion is not a standard part of our library for benchmarking, let us not introduce it for this crate.
- Kernel dispatch: I left comments about you're disptaching the kernels, please take a look.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 26 out of 27 changed files in this pull request and generated no new comments.
Suppressed comments (2)
diskann-pipnn/src/partition_kernel/tests.rs:19
PartitionTopK::leader_scalesis documented as "squared leader norms for L2" (and cosine uses unsquared norms), but this test helper feeds unsquared values for the L2 case. That makes the test data inconsistent with the public contract and can mask mistakes in distance computation. Consider squaring the L2 norms here so the tests exercise the intended inputs.
let leader_scales = match metric {
Metric::L2 => (0..leaders).map(|leader| (leader + 1) as f32).collect(),
Metric::Cosine => (0..leaders)
diskann-pipnn/src/partition_kernel.rs:252
- For the L2 path, the SIMD chunk uses
mul_add_simd(fused multiply-add) but the scalar tail usesnorm - 2.0 * dot(non-fused). This can introduce small rounding differences between SIMD and tail elements, which can change ordering/tie behavior right at SIMD-width boundaries. Usef32::mul_addfor the scalar tail so both paths compute the same value shape.
|dot, norm| F::splat(arch, -2.0).mul_add_simd(dot, norm),
|dot, norm| norm - 2.0 * dot,
@microsoft-github-policy-service agree company="Microsoft" |
PiPNN callers already bound or check their matrix shapes. Keep the shared constructor change and its tests outside this PR.
The merge with main enables clippy::allow_attributes. Use one scoped dead-code expectation until graph construction wires the kernels.
Mark Hildebrand (hildebrandmw)
left a comment
There was a problem hiding this comment.
I was hoping to get through the review today, but have run out of time. Please see the inline comments for the partial review.
One thing from a developer quality-of-life standpoint is that this PR makes heavy use of rstest to stamp out table-driven tests. As is, it increases the total number of test function in diskann from 354 to 1605! This adds tremendous noise to the test output since just about 78% of the tests are PiPNN related and the test names go on forever. E.g.
PASS [ 0.014s] diskann graph::pipnn::partition_metric::tests::point_to_leader_scores_match_scalar_distances::case_4_inner_product::point_count_2_3::leader_count_1_1::dimensions_06_15
Many of these can be easily bundled up into an internal table driven test and would greatly increase the signal to noise ratio.
There is another orthogonal concern to using rstest this way and that is that Miri + nextest has a start up time that is proportional to the number of tests. You can try this out with diskann-wide where Miri takes forever to start up due to the 1000s of tests. Test spamming like this makes the Miri experience for developers unnecessarily worse.
Please consolidate tests that are merely using rstest to run through combinations of values.
Each kernel now calls one top-k function: - select_top_k_ids ranks partition rows; select_top_k_symmetric runs the leaf pair scan. Both take k from the output width and choose fixed-size or slice nearest sets internally. - Remove Ranker, TopKVisitor, BatchRanker, Batch, BatchVisitor, with_topk, and with_batch. - Share output-row validation and checked distance scratch between the kernels. Remove leaf_neighbor_count, LeafKernelError, and UNASSIGNED_LEADER. - Load SIMD groups through one unsafe load whose bounds come from &[f32; LANES]. A load from a copied group cost 11% of leaf ranking time on AVX2. - Leaf CosineNormalized delegates to InnerProduct and adds 1. - Top-k and kernel tests run on every supported architecture and cross SIMD groups. Point the nightly Miri step at the renamed tests.
The kernel entry points already return an error for a bad output shape, so the top-k functions now check their shape preconditions with debug assertions only. Select distance rows by index. A distance matrix without columns then leaves every slot UNASSIGNED instead of panicking in row_iter.
|
Pushed b0ccf73 and cc6da38, and merged them into #1290, #1291, #1294, and #1295.
|
Mark Hildebrand (hildebrandmw)
left a comment
There was a problem hiding this comment.
Thanks - this has been quite a journey! A few small last comments on my end mostly about protecting ourself against future changes of the SIMD width assumption.
Other than that, happy to see this merge!
The unsafe group load reads A::Vector::LANES values from a &[f32; LANES] group. Assert that the two counts match, so a future per-architecture width fails loudly. The per-architecture top-k tests take their group boundaries from A::Vector::LANES.
Summary
Add GEMM-backed distance computation and SIMD top-k selection for PiPNN behind the opt-in diskann/pipnn feature.
The numerical layer supports two operations:
This is part 2 of the PiPNN stack. Recursive graph construction integrates these kernels in #1290.
Design
Separate metric computation from neighbor selection. Metric implementations handle L2, cosine, normalized cosine, and inner product. Partition leaders retain metric-specific norms for reuse across point stripes. Kernels choose which distance slices
to process; the shared TopK implementation handles SIMD loading, threshold filtering, scalar tails, and sorted insertion.
Reuse each leaf distance for both endpoints. New diskann-linalg helpers compute or accumulate the lower triangle of a Gram matrix using faer. Leaf selection scans the strict lower triangle, excludes self-neighbors, and shares each loaded distance
block between both endpoint updates.
Specialize common top-k capacities. Capacities 1–10 use compile-time specialization to enable insertion-loop unrolling; larger capacities use the runtime path. Typical partition fanouts are 10 and 3, and typical leaf-neighbor counts are 2 or 3.
Architecture-specific selection runs through diskann-wide at the TopK entry points.
Keep storage reusable. TopK borrows caller-owned candidate buffers. Kernels manage distance and ranking scratch without exposing SIMD block types to their callers.
The linear-algebra helpers validate matrix shapes and size-product overflow.