Skip to content

perf: optimize phrase match scorer - #30

Merged
zhengbuqian merged 3 commits into
mainfrom
optimize-phrase-match-scorer
Sep 21, 2026
Merged

zhengbuqian merged 3 commits into
mainfrom
optimize-phrase-match-scorer

Conversation

@zhengbuqian

Copy link
Copy Markdown
Collaborator

Summary

Optimize Tantivy's no-scoring phrase-existence path without changing query semantics:

  • Two-term slop existence uses a two-pointer check instead of materializing PositionSpans.
  • Span buffers are allocated only for slop queries with more than two terms.
  • Multi-term span matching stops scanning next_positions once a position inside the current span is found.

Benchmark

Standalone 1M-document in-memory index, seed 42, Count collector (no scoring), 10 warmup / 50 measured iterations. Matrix:

  • Hit rates: 10% / 30% / 50% / 90%
  • Phrase lengths: 2 / 3 / 5 / 10 / 30 terms
  • Slop: 0 / 1 / 2

Average median speedup across the four hit rates:

Terms Slop 0 Slop 1 Slop 2
2 1.77% 2.98% 1.23%
3 0.94% 13.55% 14.30%
5 3.70% 21.78% 22.14%
10 2.15% 26.95% 26.99%
30 2.69% 29.24% 28.91%

Across all 40 slop>0 cases the average median speedup is 18.81%. Exact-phrase (slop 0) averages 2.25% and is treated as run-to-run noise because that path was not algorithmically changed. Gains grow with phrase length on slop 1/2, peaking around 28–31% for 30-term queries.

Benchmark sources were used locally for measurement and are not included in this PR.

Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
@zhengbuqian
zhengbuqian force-pushed the optimize-phrase-match-scorer branch from 7a0809d to 66288ab Compare September 21, 2026 07:45
Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
}
best_match.clear();
best_match.push(*prev_qualified);
break;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This early exit suggests a stronger simplification for the no-scoring existence path.

For a fixed span [L, R] and sorted next_positions, the expanded width is:

R - p,  p < L
R - L,  L <= p <= R
p - L,  p > R

Therefore the only non-dominated transitions are:

  1. any position inside [L, R], which keeps the span unchanged; or
  2. if none exists, the predecessor max(p < L) and successor min(p > R).

All positions farther left or right are dominated and do not need to be scanned. This permits a dedicated phrase_exists frontier algorithm:

  • initialize [p, p] for the first term positions;
  • sweep the ordered span frontier and the next ordered position list with one monotonic cursor;
  • emit at most the unchanged span or the nearest left/right expansions whose widths fit max_slop;
  • remove duplicates and any interval that contains a smaller interval.

After pruning, both span endpoints are strictly increasing. Each position is advanced at most once and each candidate is pushed/popped at most once, so one intermediate-term merge becomes O(F + P) for F frontier spans and P positions. The current nested scan remains O(S * P) in the worst case because right-side positions can be revisited for multiple spans. For K terms with roughly P states per round, this changes the existence path from O(K * P^2) to O(K * P).

Both nearest-side candidates must be retained when valid; choosing only the smaller current width is not sufficient. For example, with max_slop = 90, current span [10, 15], next positions [9, 100], and a following term at [100], keeping only [9, 15] later produces width 91 and misses the valid path [10, 100], whose final width remains 90.

This optimization should be isolated to phrase_exists: containment pruning is safe for boolean existence, while phrase_count may need occurrence-chain information.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point, you can implement this in a separate pr.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One missing detail is how to generate this frontier in stack order without needing middle insertion.

Let the input frontier be I_i = [L_i, R_i], with both L_i and R_i strictly increasing after containment pruning, and let j_i = lower_bound(positions, L_i). Since L_i increases, j_i is non-decreasing.

For each span:

  • If positions[j_i] <= R_i, emit only [L_i, R_i]; an in-range position leaves the span unchanged and therefore dominates every exterior expansion.
  • Otherwise, the only candidates are [positions[j_i - 1], R_i] and [L_i, positions[j_i]], when those positions exist and the width fits max_slop.
  • If consecutive spans have the same j_i, emit the left candidate only for the first such span. All later left candidates have the same left endpoint and a larger R_i, so the first one is contained in and dominates them. The right candidates have the same right endpoint and increasing left endpoints, so each later one replaces the previous one.

This local rule makes the surviving candidates arrive with non-decreasing left endpoints. While j_i is unchanged, the repeated left candidate is removed and all remaining candidates use the increasing L_i. When j_i advances, the new predecessor is at least the previous successor, which is at least the preceding span left endpoint, so the candidate left endpoint cannot move backward.

The output can then be maintained with a monotonic stack whose L and R are both strictly increasing:

append(candidate):
    if top.L == candidate.L and top.R <= candidate.R:
        discard candidate
    else:
        while stack is not empty and top.R >= candidate.R:
            pop top
        push candidate

The proof that comparing only the top is sufficient is now simple. Candidates arrive in non-decreasing L, so a new candidate can never belong before a retained span by L. If its R is not greater than the top R, it is contained in the top and dominates it. Because stack R values are increasing, every other interval dominated by the candidate is a contiguous suffix, removed by the same loop. Once top.R < candidate.R, neither interval contains the other, so the candidate belongs exactly at the end. Therefore no insertion into the middle is required.

A raw left candidate can have both endpoints below the current top, but only when its predecessor cursor did not advance; that candidate has already been dominated by the first left candidate in the same cursor run and must be discarded before stack insertion.

j_i advances at most P times, and every candidate is pushed and popped at most once, so a merge round is O(F + P) time and O(F_next) output space.

@aoiasd

aoiasd commented Sep 21, 2026

Copy link
Copy Markdown

/lgtm

@zhengbuqian
zhengbuqian merged commit 9ce00de into main Sep 21, 2026
5 of 7 checks passed
@zhengbuqian
zhengbuqian deleted the optimize-phrase-match-scorer branch September 21, 2026 11:51
zhengbuqian pushed a commit that referenced this pull request Oct 8, 2026
## Summary

Follow up on #30 with a dedicated algorithm for no-scoring phrase
existence queries with more than two terms and non-zero slop.

- Keep both nearest-left and nearest-right expansions when neither
contains the other.
- Prune spans only through containment, which is safe for boolean
existence.
- Maintain the minimal span frontier with one position cursor and a
monotonic stack.
- Leave `phrase_count` semantics unchanged and preserve dedicated
exact-phrase and two-term paths.

This also fixes a false negative in the previous minimum-current-width
pruning. With adjusted positions `[10], [15], [9, 100], [100]` and slop
90, `[9, 15]` is narrower locally but cannot match the final 100; the
non-dominated `[10, 100]` branch must also be retained.

## Complexity

For `F` current frontier spans and `P` next positions, each position
advances once and each candidate is pushed and popped at most once:

- previous worst case: `O(F * P)` per intermediate term
- new worst case: `O(F + P)` per intermediate term
- space: `O(F_next)` using the existing reusable span buffers

The source comments document why candidates arrive with non-decreasing
left endpoints and why containment removals are always a stack suffix.

## Benchmark

### Isolated `phrase_exists` benchmark

This benchmark isolates the `phrase_exists` operation. Each iteration
receives pre-generated sorted positions for all terms in one document
and runs phrase matching only: initial span creation, intermediate
frontier merges, and the final-term existence check. It excludes index
construction, postings decoding, document intersection, collectors, and
scoring.

- Baseline: the implementation before this PR, from #30 merge commit
`9ce00de9a822b2a1cfa8d2b91454d7ae017969b9`
- Command: `taskset -c 6 target/release/algorithm 2048 11 100`
- One fixed CPU core
- 2,048 documents per cell with a 50% hit rate
- 11 samples per cell; each sample is calibrated to run for at least 100
ms; report the median
- Phrase lengths: 3 / 5 / 10 / 30 terms
- Positions per term: 1 / 4 / 16 / 64
- Slop: 1 / 4 / 16 / 64
- Measurement order rotates between implementations
- Every timed workload asserts identical hit counts between the baseline
and optimized implementations

The following values are nanoseconds per document. Each row averages the
medians from its four slop values. Speedup is always calculated against
the fixed baseline above.

| Terms | Positions / term | Baseline | Optimized | Speedup |
|------:|-----------------:|---------:|----------:|--------:|
| 3 | 1 | 9.73 | 7.79 | **20.00%** |
| 3 | 4 | 68.11 | 47.56 | **30.15%** |
| 3 | 16 | 260.02 | 193.55 | **25.28%** |
| 3 | 64 | 1168.63 | 768.73 | **31.34%** |
| 5 | 1 | 15.58 | 11.46 | **26.45%** |
| 5 | 4 | 118.58 | 70.79 | **40.49%** |
| 5 | 16 | 371.79 | 258.11 | **30.65%** |
| 5 | 64 | 1938.25 | 1247.55 | **31.34%** |
| 10 | 1 | 30.97 | 20.25 | **34.61%** |
| 10 | 4 | 239.89 | 118.26 | **51.09%** |
| 10 | 16 | 538.04 | 320.51 | **41.04%** |
| 10 | 64 | 2615.11 | 1618.27 | **35.12%** |
| 30 | 1 | 104.02 | 62.59 | **39.73%** |
| 30 | 4 | 687.37 | 268.51 | **61.37%** |
| 30 | 16 | 1177.39 | 537.01 | **55.19%** |
| 30 | 64 | 3900.96 | 2337.06 | **38.41%** |

Speedup for each slop value:

| Terms | Positions / term | Slop 1 | Slop 4 | Slop 16 | Slop 64 |
|------:|-----------------:|-------:|-------:|--------:|--------:|
| 3 | 1 | 20.61% | 20.06% | 19.69% | 19.62% |
| 3 | 4 | 29.53% | 30.37% | 29.70% | 30.99% |
| 3 | 16 | 23.48% | 23.87% | 24.49% | 29.28% |
| 3 | 64 | 23.33% | 24.86% | 29.25% | 47.92% |
| 5 | 1 | 28.23% | 24.96% | 25.37% | 27.22% |
| 5 | 4 | 48.47% | 38.25% | 37.82% | 37.42% |
| 5 | 16 | 32.50% | 30.44% | 29.02% | 30.65% |
| 5 | 64 | 26.51% | 26.08% | 28.39% | 44.39% |
| 10 | 1 | 32.93% | 33.47% | 36.25% | 35.78% |
| 10 | 4 | 63.31% | 56.49% | 43.00% | 41.57% |
| 10 | 16 | 44.26% | 43.27% | 39.58% | 37.06% |
| 10 | 64 | 32.86% | 32.86% | 31.96% | 42.82% |
| 30 | 1 | 33.06% | 38.30% | 38.45% | 49.12% |
| 30 | 4 | 71.29% | 68.97% | 57.36% | 47.86% |
| 30 | 16 | 60.00% | 58.55% | 53.56% | 48.65% |
| 30 | 64 | 36.08% | 36.27% | 38.15% | 43.16% |

Average speedup by positions per term:

| Positions / term | Speedup |
|-----------------:|--------:|
| 1 | **30.20%** |
| 4 | **45.78%** |
| 16 | **38.04%** |
| 64 | **34.06%** |

Average speedup by slop:

| Slop | Speedup |
|-----:|--------:|
| 1 | **37.90%** |
| 4 | **36.69%** |
| 16 | **35.13%** |
| 64 | **38.34%** |

The unweighted average across all 64 cells is **37.02%**. Every measured
cell improved: the range was **19.62%** to **71.29%**. Three full runs
differed by no more than 0.42 percentage points in any
positions-per-term bucket.

The generator rejected and regenerated three candidate documents on
which the legacy implementation produced a different result, so semantic
differences did not affect the timing comparison. Benchmark sources were
used locally and are not included in the PR.

### End-to-end phrase-query benchmark

This benchmark measures the complete no-scoring query path through
postings decoding, document intersection, phrase matching, and the
`Count` collector. Index construction is excluded.

- Baseline: the same #30 merge commit,
`9ce00de9a822b2a1cfa8d2b91454d7ae017969b9`
- Command: `taskset -c 6 <binary> <version> 1024 11 100`
- B–O–B execution order; each cell compares the optimized median with
the average of the two baseline medians
- 1,024 documents per cell; every document contains every query term and
50% satisfy the phrase
- 11 samples per cell; each sample is calibrated to run for at least 100
ms; report the median
- Text lengths: 512 / 2,048 / 8,192 tokens
- Term repetitions per document: 1 / 4 / 16
- Phrase lengths: 2 / 3 / 5 / 10 / 30 terms
- Slop: 4

Speedup averaged over the three text lengths:

| Term repetitions | 2 terms | 3 terms | 5 terms | 10 terms | 30 terms |
|-----------------:|--------:|--------:|--------:|---------:|---------:|
| 1 | **7.32%** | **24.74%** | **23.87%** | **18.45%** | **13.66%** |
| 4 | **4.66%** | **29.36%** | **34.02%** | **40.47%** | **47.27%** |
| 16 | **2.61%** | **40.84%** | **52.48%** | **59.92%** | **64.05%** |

For the changed path (3 or more terms), the average speedup across all
36 cells was **37.43%**, ranging from **11.94%** to **64.57%**; every
cell improved. Average speedup was stable across text lengths:
**37.98%** at 512 tokens, **37.89%** at 2,048, and **36.41%** at 8,192.

The two-term control path had no measured regression: all 9 cells
improved, with an average of **4.86%** and a minimum of **1.72%**. The
two baseline passes differed by **1.42%** on average and at most
**5.18%**, so the smaller control-path gains should be treated as
benchmark noise rather than attributed to this PR.

An additional 200 ms B–O–B regression guard compared the final function
split with the preceding version of this PR. The final version was
**1.74%** faster overall and **2.47%** faster across the changed
3-or-more-term path. The two-term control measured **-1.19%**, below the
two baseline passes' **2.60%** mean absolute difference, so no
systematic regression was observed.

## Verification

- End-to-end no-scoring regression through the `Count` collector
- Duplicate-position regression through a compound-word tokenizer and
the `Count` collector
- Exhaustive comparison against a brute-force reference for all
four-term non-empty position sets over `0..4` and slop `0..=3`
- 1,000-case randomized comparison against brute force for 3–10 terms
and 1–5 positions per term
- Strict frontier invariant checked after every intermediate merge
- `cargo test query::phrase_query --lib` — 32 passed, 1 ignored
- `cargo test --lib` — 911 passed, 5 ignored
- `cargo build --all-features`
- `cargo clippy --tests`
- `cargo +nightly fmt --all -- --check`
- `cargo +nightly bench --no-run --profile=dev --all-features`

Closes #32

Signed-off-by: aoiasd <zhicheng.yue@zilliz.com>
zhengbuqian pushed a commit that referenced this pull request Oct 10, 2026
## Summary

Follow up on #33: prevent a no-scoring `PhraseQuery` from using the same
raw token position for multiple occurrences of the same query term.

For example, `a a` with slop 1 must not match a document containing only
one `a`. The previous adjusted-position matcher can reuse that
occurrence at both query offsets and incorrectly report a match.

- Detect repeated terms once when constructing the immutable query;
share occurrence IDs through `Arc` across query clones and segment
weights.
- Dispatch only repeated-term, scoring-disabled queries to a dedicated
weight/scorer. Keep the existing schema/position validation.
- Build a minimal span frontier for each distinct term, consuming
distinct raw positions for its repeated occurrences, then merge the
frontiers.
- Reuse position, cursor, and span buffers across documents. Use
monotone position cursors rather than repeated binary searches, with a
direct `n == k` path.
- Leave the existing unique-term scorer and its exact-phrase/two-term
fast paths unchanged.

## Algorithm and correctness

Adjusted positions remain `raw_position + max_query_offset -
query_offset`; a match requires their maximum minus minimum to be at
most slop. This PR does not change that slop definition.

For one term appearing at `k` query offsets:

1. Deduplicate its sorted raw positions and sort its normalized offsets
in descending order. Fewer than `k` distinct positions means no match.
2. It is sufficient to assign increasing raw positions to descending
offsets. Uncrossing a reversed pair cannot enlarge the adjusted-position
hull: the uncrossed adjusted positions both lie inside the crossed hull.
Thus an increasing assignment dominates any reversed assignment.
3. For an adjusted-position lower bound `B`, greedily choose the
earliest eligible unused position for each offset. Any other increasing
assignment satisfying `B` chooses componentwise later positions, so this
greedy assignment minimizes the right endpoint.
4. Emit its hull if it satisfies slop, then raise `B` past its left
endpoint. Among assignments satisfying the current bound `B`, any
assignment with no larger left endpoint has no smaller right endpoint
and is dominated; assignments with larger left endpoints are considered
by subsequent rounds.
5. Each offset's lower bound and its predecessor's chosen index only
increase. Its cursor therefore never moves backward, including across
candidate rounds.

Merge distinct-term frontiers using their strictly increasing left/right
endpoints. For each current span, advance to the last next-term span
with `right <= current.right`: retain the current span on containment,
otherwise consider the best left-only expansion, subsequent double-sided
expansions, and the first right-only expansion. Consumed double-sided
spans cannot improve a later current span; retain the right-only
candidate for possible reuse.

Candidates arrive with non-decreasing left endpoints. Equal-left
candidates retain the smaller right endpoint; a smaller/equal right
endpoint removes a contained suffix from the output stack. Each
candidate is pushed/popped at most once. Containment pruning is safe
because choices from different term groups cannot compete for the same
term occurrence. Width filtering is safe because later merges cannot
shrink a hull.

For `n` raw positions, `k` occurrences, and `F` evaluated candidate
rounds, the repeated-term builder takes `O(k*n + F*k)` time and `O(k)`
reusable cursor space, plus the output frontier. Here `F` includes
slop-rejected rounds, not only emitted spans. The `n == k` case takes
`O(k)`. Merging frontiers of lengths `A` and `B` takes `O(A+B)` time,
plus output storage.

## Verification

- Indexed no-score regressions for `a a`, `a a a`, `a b a`, and exact `a
a b`, using the `Count` collector.
- Query-cache test for custom offset ordering, clone sharing, and
changes to slop.
- No-score `explain` regression for matching documents, earlier/later
non-matches, absent terms, and an exhausted scorer. Reject a document
before seeking if the new scorer has already passed it.
- 2,646 exhaustive repeated-term-frontier comparisons against all
injective assignments.
- Three randomized properties: repeated-term frontier versus brute
force, frontier merge versus the Cartesian product, and complete scorer
results versus an independent injective-assignment oracle. The complete
scorer property covers duplicate raw positions, repeated/custom query
offsets, multiple documents, adjacent slop values, `advance`, `seek`,
and reused buffers.
- Extended property runs: 20,000 cases per property in each of three
runs (one random seed and fixed seeds `20261009` and `42`), with no
mismatches.
- Separate real-index verification: 8,000 randomized queries over two
96-document corpora, comparing actual matching document sets against an
independent brute-force oracle. All 768,000 query/document decisions
agreed, including 482,244 negative decisions. This external verifier is
not included in the PR.
- `cargo test --lib`: 918 passed, 5 ignored before the final
explain-only guard; the additional explain regression passed separately.
- `cargo test --lib --features
mmap,stopwords,lz4-compression,zstd-compression,failpoints,quickwit --
--test-threads=1`: 919 passed, 5 ignored including the final guard.
- An earlier parallel full-suite rerun had failures in unchanged indexer
tests `create_index_test_max_merge_issue_1035` and
`test_delete_during_merge`; both passed individually on rerun. The
earlier pre-guard parallel feature run passed all 918 tests. These
intermittent failures are disclosed rather than treated as a fully clean
parallel-test history.
- `cargo build --all-features`, `cargo clippy --lib --tests`, `cargo
+nightly fmt --all -- --check`, and `git diff --check`: passed. Existing
unrelated warnings remain.

## End-to-end benchmark

Benchmark sources/artifacts are local only and are not included in this
PR. Measurements cover postings decoding, document intersection, phrase
matching, and `Count`, excluding index/query construction. Queries are
reused, so this does not measure the new one-time duplicate-detection
cost.

The timings predate the final explain-only guard and its test. The
matching algorithm is unchanged; these are not a fresh benchmark of the
final commit's binary layout.

- Intel i7-8700, rustc 1.92.0, release build, pinned to CPU 6; the SMT
sibling is not isolated and frequency is not fixed.
- 512 documents in a single segment; every timed document genuinely
matches under both implementations. Every search asserts 512 hits.
- Text lengths: 512 / 2,048 / 8,192 tokens; repetitions of a 30-token
block: 1 / 4 / 16; query lengths: 2 / 3 / 5 / 10 / 30.
- Query patterns: all terms unique; adjacent pairs repeated; all query
terms identical. The corresponding positions per distinct term are `R`,
`2R`, and `30R` respectively, where `R` is block repetitions.
- Slop: 0 / 1 / 4 / 16 / 64 at length 2,048, and 4 at the other lengths.
Total: 315 cells, 105 per pattern.
- Seven samples targeting 50 ms each; median per cell. The initial
repair was measured before and after the final cursor implementation;
use the geometric mean of the two initial-repair medians for that
comparison.
- Historical baseline stays fixed at the implementation before #33,
namely #30 merge commit `9ce00de9a822b2a1cfa8d2b91454d7ae017969b9`. The
existing #33 implementation is also shown separately; its reference
source tree matches the parent of this PR,
`596a0f65ba24dd94b15ef97b9b9a909b610b8f22`. Historical baseline
measurements were retained, not regenerated.

Geometric-mean time ratios, normalized per cell to the fixed historical
baseline (`1.000`; lower is better):

| Query pattern | Existing #33 | Initial repair with binary searches |
Final repair | Final / #33 |
|---|---:|---:|---:|---:|
| Unique | 0.631 | 0.660 | 0.648 | 1.027 |
| Adjacent repeated pairs | 0.533 | 1.014 | 0.581 | 1.089 |
| All terms identical | 0.407 | 2.087 | 0.353 | 0.868 |

The monotone-cursor implementation reduces time by 42.70% for repeated
pairs and 83.09% for identical terms versus the initial repair. All 210
repeated-term cells improve versus that initial repair. These are not
universal speedups versus #33: repeated pairs remain 8.90% slower on
aggregate, while identical terms are 13.21% faster on aggregate. The
unique-term control is 2.70% slower on aggregate versus #33.

Slop breakdown at text length 2,048 (15 cells per pattern/slop), final
time / existing #33 time:

| Slop | Unique | Repeated pairs | Identical terms |
|---:|---:|---:|---:|
| 0 | 1.026 | 1.186 | 1.186 |
| 1 | 1.023 | 1.065 | 0.827 |
| 4 | 1.023 | 1.070 | 0.828 |
| 16 | 1.029 | 1.071 | 0.829 |
| 64 | 1.028 | 1.073 | 0.826 |

A high-density, short repeated query remains expensive: identical terms,
length 2,048, `R=16` (480 positions), 2 query terms, slop 4 costs
384.446 us in the fixed historical baseline, 379.753 us in #33,
9,631.178 us in the initial repair, and 1,530.291 us in the final
repair. The final repair is still 4.03x #33 here. It constructs the
complete valid repeated-term frontier rather than stopping after the
first complete match.

Separate collision fixtures are excluded from performance aggregates
because the old matcher returns different results. Across 50 fixtures,
the final repair returns the expected 256 matches; both legacy
implementations incorrectly return 512 in 40 fixtures.

### Performance limitations and diagnosis

No claim of regression-free ordinary queries is made. The
final/initial-repair comparison includes seven unique-term cells slower
by more than 5%, with a maximum of 15.75%. Follow-up CPU-pinned A-B-B-A
runs with 11 samples targeting 300 ms reproduced ordinary-query search
regressions, including 10.55% and 7.23%, while corresponding same-binary
A/A search runs differed by about 0–2.3%.

For the 10.55% case, weight creation got slightly faster; scorer
construction increased by about 53 ns, whereas the full search increased
by about 3.76 us. The increase is mainly in document scanning, not a new
per-document duplicate check. The ordinary phrase scorer source is
unchanged, and normalized assembly for its matching/advance/seek
functions is identical across the diagnostic builds.

A diagnostic-only 64-byte function-alignment build changed those two
regressions to -1.18% and -1.68%, showing layout sensitivity, but
another case still regressed 5.63%. This is evidence, not a complete
CPU-level explanation or a production fix. Hardware performance counters
were unavailable. The PR does not change compiler alignment options.

## Scope and remaining work

- This fixes scoring-disabled literal `PhraseQuery` only. Scored
matching / `phrase_count` remain on the existing matcher and can still
reuse repeated-term positions. Regex and prefix phrase queries are
unchanged and are not covered by this collision verification.
- Physical postings/intersection inputs still include each query
occurrence. Only position matching is grouped by distinct term.
- Streaming first-match termination and slop-aware skipping of invalid
candidate bounds are possible follow-ups; they are not implemented here.
- No Milvus build/live end-to-end test is claimed, and the full
workspace/doctest/bench CI matrix has not been rerun locally for this
change. Position arithmetic retains existing `u32` overflow limits.

Signed-off-by: aoiasd <zhicheng.yue@zilliz.com>
sre-ci-robot pushed a commit to milvus-io/milvus that referenced this pull request Oct 10, 2026
Advance the locked Tantivy 0.23 revision to
`fef3c2a5748bdb2b2c2f8ca1d9fc125a0d1e8b71`, including the no-scoring
phrase optimizations from
[#30](zilliztech/tantivy#30),
[#33](zilliztech/tantivy#33), and the
repeated-term correctness fix from
[#34](zilliztech/tantivy#34). For example, `a a`
with slop 1 must not match a document containing only one `a`. Milvus's
phrase-match collectors disable scoring and reach the corrected path.

### Revision and scope
- Verified #34 merged at 2026-10-10 03:14:27 UTC. Its merge commit is
the selected revision and was the upstream default branch `main` HEAD
when this update began.
- Its sole parent is the previous PR pin
`596a0f65ba24dd94b15ef97b9b9a909b610b8f22` (#33), whose parent is
`9ce00de9a822b2a1cfa8d2b91454d7ae017969b9` (#30). GitHub state and local
Git ancestry agree.
- This follow-up changes only nine Tantivy workspace source hashes in
`internal/core/thirdparty/tantivy/tantivy-binding/Cargo.lock` (+9/-9).
The upstream increment changes four phrase-query source files, with no
manifest changes.
- The complete PR remains lockfile-only (+20/-11 from the branch base),
retaining the original rust-stemmers 1.3.0 addition required by modern
Tantivy. Legacy Tantivy 0.21.1, its rust-stemmers 1.2.0 pin, all
registry packages, and Cargo.toml remain unchanged.

### Current local validation (2026-10-10, fef3c2a57)
- Normal retained `build-server.sh`: passed; release Rust bindings,
Milvus core, C++ test targets and Go server rebuilt with dependency
checks. FFI header and artifact/shared-library checks passed.
- Binding release text/phrase tests: 4 passed. Analyzer tests: 6 passed.
- Upstream `cargo +1.89 test --locked --lib`: 919 passed, 5 ignored.
Phrase subset: 39 passed, 1 ignored. Includes repeated-term indexed
queries, exhaustive and randomized injective-assignment regressions,
clone sharing, and explain tests.
- Freshly linked C++ `TextMatch.*:PackedTextMatchAsyncLoadTest.*`: 33
passed, 2 failed because existing LOB fixtures write
`/sealed_text_lob_index_*` and `/growing_text_lob_index_*` on a
read-only filesystem. No fixture edits were applied this round; no
#53986 changes are included.
- `git diff --check` and parsed lockfile comparison: passed; only the
nine source hashes changed in this follow-up.

The earlier validation of 596a0f65 is historical: its two LOB fixture
tests passed with temporary #53986 path adjustments, subsequently
reverted. That result is not claimed as a test of this revision. No
live-cluster E2E or local performance benchmark was run. Upstream
benchmark sources are not checked in; upstream #34 reports some
regressions as well as improvements, so this update makes no universal
speedup or measured Milvus performance claim. The fix covers unscored
literal PhraseQuery; scored, regex and prefix phrase semantics are
outside its scope.

---------

Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
sre-ci-robot pushed a commit to milvus-io/milvus that referenced this pull request Oct 10, 2026
)

pr: #54023

Advance the locked Tantivy 0.23 revision to
`fef3c2a5748bdb2b2c2f8ca1d9fc125a0d1e8b71`, including the no-scoring
phrase optimizations from
[#30](zilliztech/tantivy#30),
[#33](zilliztech/tantivy#33), and the
repeated-term correctness fix from
[#34](zilliztech/tantivy#34). For example, `a a`
with slop 1 must not match a document containing only one `a`. Milvus's
phrase-match collectors disable scoring and reach the corrected path.

### Revision and scope
- Verified #34 merged at 2026-10-10 03:14:27 UTC. Its merge commit is
the selected revision and was the upstream default branch `main` HEAD
when this update began.
- Its sole parent is the previous PR pin
`596a0f65ba24dd94b15ef97b9b9a909b610b8f22` (#33), whose parent is
`9ce00de9a822b2a1cfa8d2b91454d7ae017969b9` (#30). GitHub state and local
Git ancestry agree.
- This follow-up changes only nine Tantivy workspace source hashes in
`internal/core/thirdparty/tantivy/tantivy-binding/Cargo.lock` (+9/-9).
The upstream increment changes four phrase-query source files, with no
manifest changes.
- The complete PR remains lockfile-only (+20/-11 from the branch base),
retaining the original rust-stemmers 1.3.0 addition required by modern
Tantivy. Legacy Tantivy 0.21.1, its rust-stemmers 1.2.0 pin, all
registry packages, and Cargo.toml remain unchanged.

### Current local validation (2026-10-10, fef3c2a57)
Used the separate 3.0 worktree and target directory, with the branch's
default features and release profile; no master or 2.6 build outputs
were reused.
- `cargo +1.89 build --release --locked`: passed. Generated files
unchanged.
- Binding release text/phrase tests: 4 passed. Analyzer tests: 6 passed.
- Shared selected upstream revision tested this round: 919 library tests
passed, 5 ignored; 39 phrase tests passed, 1 ignored. Includes #34's
repeated-term indexed queries, exhaustive/randomized property tests,
clone sharing and explain regressions.
- `git diff --check` and parsed lockfile comparison: passed; only nine
source hashes changed in this follow-up.

Earlier results at 596a0f65 are historical. No full release-branch
C++/Go build, live-cluster E2E or local performance benchmark was run.
No test/source adjustments or #53986 changes are included. Upstream #34
reports some performance regressions; no universal speedup or measured
Milvus gain is claimed. The fix covers unscored literal PhraseQuery, not
scored, regex or prefix phrases.

---------

Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
sre-ci-robot pushed a commit to milvus-io/milvus that referenced this pull request Oct 10, 2026
)

pr: #54023

Advance the locked Tantivy 0.23 revision to
`fef3c2a5748bdb2b2c2f8ca1d9fc125a0d1e8b71`, including the no-scoring
phrase optimizations from
[#30](zilliztech/tantivy#30),
[#33](zilliztech/tantivy#33), and the
repeated-term correctness fix from
[#34](zilliztech/tantivy#34). For example, `a a`
with slop 1 must not match a document containing only one `a`. Milvus's
phrase-match collectors disable scoring and reach the corrected path.

### Revision and scope
- Verified #34 merged at 2026-10-10 03:14:27 UTC. Its merge commit is
the selected revision and was the upstream default branch `main` HEAD
when this update began.
- Its sole parent is the previous PR pin
`596a0f65ba24dd94b15ef97b9b9a909b610b8f22` (#33), whose parent is
`9ce00de9a822b2a1cfa8d2b91454d7ae017969b9` (#30). GitHub state and local
Git ancestry agree.
- This follow-up changes only nine Tantivy workspace source hashes in
`internal/core/thirdparty/tantivy/tantivy-binding/Cargo.lock` (+9/-9).
The upstream increment changes four phrase-query source files, with no
manifest changes.
- The complete PR remains lockfile-only (+20/-11 from the branch base),
retaining the original rust-stemmers 1.3.0 addition required by modern
Tantivy. Legacy Tantivy 0.21.1, its rust-stemmers 1.2.0 pin, all
registry packages, and Cargo.toml remain unchanged.

### Current local validation (2026-10-10, fef3c2a57)
Used the separate 2.6 worktree and target directory, with the branch's
default features and release profile; no master or 3.0 build outputs
were reused.
- `cargo +1.89 build --release --locked`: passed. Generated files
unchanged.
- Binding release text/phrase tests: 3 passed.
- Analyzer tests: 4 passed, 1 failed. The unchanged legacy
`index_writer_v5::analyzer::analyzer::tests::test_lindera_analyzer`
requests IPADIC while the default feature set enables no dictionary,
yielding `lindera tokenizer with invalid dict_kind`. This reproduces the
previously documented configuration limitation; legacy code and pins are
unchanged.
- Shared selected upstream revision tested this round: 919 library tests
passed, 5 ignored; 39 phrase tests passed, 1 ignored, including #34's
repeated-term, property, clone-cache and explain regressions.
- `git diff --check` and parsed lockfile comparison: passed; only nine
source hashes changed in this follow-up.

Earlier results at 596a0f65 are historical. No full release-branch
C++/Go build, live-cluster E2E or local performance benchmark was run.
No test/source adjustments or #53986 changes are included. Upstream #34
reports some performance regressions; no universal speedup or measured
Milvus gain is claimed. The fix covers unscored literal PhraseQuery, not
scored, regex or prefix phrases.

---------

Signed-off-by: Buqian Zheng <zhengbuqian@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants