Summary
In the vector-based span disambiguation feature, the context window is hashed using binary token presence rather than within-window term frequency. A token that appears multiple times in a single window contributes the same as a token that appears once.
This is a deliberate modeling decision, not a bug, and there is a standing TODO in the code asking whether it is the right one. This issue tracks evaluating it.
Where
VectorBasedSpanDisambiguationService.hash():
// We're only looking for what the window has. How many of each token is irrelevant.
// TODO: But is it irrelevant though? If a word occurs more often than others
// it is probably more indicative of the type than a word that only occurs once.
vector[hash] = 1;
So a window ["phone", "call", "phone"] produces the same vector as ["phone", "call"].
Important scope note
This is only about repetition within a single window. Repetition across documents already accumulates as counts in the per-filter-type vectors (see InMemoryVectorService / FileBasedVectorService), and that corpus-level weighting is already exercised by tests. The open question is narrowly whether a token repeated inside one short window should get extra weight.
Trade-offs
For counting within-window frequency (vector[hash] += 1):
- A word repeated in the immediate vicinity is plausibly more indicative of the type (the TODO's intuition).
- Moves from binary bag-of-words toward term-frequency weighting.
For keeping binary presence (status quo):
- Windows are short (a handful of tokens per side), so within-window repeats are rare and often incidental.
- Counting them risks letting one repeated common word dominate a single training example.
- The ambiguous span's query vector is built by the same
hash(), so switching to counts changes both the learned side and the query side; the interaction with cosine normalization is non-obvious and should be measured, not assumed.
Why low priority
The impact is small and the direction of the effect is genuinely uncertain — the change could help or hurt accuracy depending on the data. This cannot be settled by a unit test; it needs a labeled evaluation set / accuracy benchmark. Until such an evaluation exists, the status quo (binary presence) is a reasonable default.
Suggested next step
Defer until there is a disambiguation accuracy evaluation harness, then A/B the two approaches (binary presence vs. within-window term frequency) on labeled data and decide based on measured accuracy.
Summary
In the vector-based span disambiguation feature, the context window is hashed using binary token presence rather than within-window term frequency. A token that appears multiple times in a single window contributes the same as a token that appears once.
This is a deliberate modeling decision, not a bug, and there is a standing
TODOin the code asking whether it is the right one. This issue tracks evaluating it.Where
VectorBasedSpanDisambiguationService.hash():So a window
["phone", "call", "phone"]produces the same vector as["phone", "call"].Important scope note
This is only about repetition within a single window. Repetition across documents already accumulates as counts in the per-filter-type vectors (see
InMemoryVectorService/FileBasedVectorService), and that corpus-level weighting is already exercised by tests. The open question is narrowly whether a token repeated inside one short window should get extra weight.Trade-offs
For counting within-window frequency (
vector[hash] += 1):For keeping binary presence (status quo):
hash(), so switching to counts changes both the learned side and the query side; the interaction with cosine normalization is non-obvious and should be measured, not assumed.Why low priority
The impact is small and the direction of the effect is genuinely uncertain — the change could help or hurt accuracy depending on the data. This cannot be settled by a unit test; it needs a labeled evaluation set / accuracy benchmark. Until such an evaluation exists, the status quo (binary presence) is a reasonable default.
Suggested next step
Defer until there is a disambiguation accuracy evaluation harness, then A/B the two approaches (binary presence vs. within-window term frequency) on labeled data and decide based on measured accuracy.