Skip to content

Evaluate within-window token frequency weighting in span disambiguation #312

Description

@jzonthemtn

Summary

In the vector-based span disambiguation feature, the context window is hashed using binary token presence rather than within-window term frequency. A token that appears multiple times in a single window contributes the same as a token that appears once.

This is a deliberate modeling decision, not a bug, and there is a standing TODO in the code asking whether it is the right one. This issue tracks evaluating it.

Where

VectorBasedSpanDisambiguationService.hash():

// We're only looking for what the window has. How many of each token is irrelevant.
// TODO: But is it irrelevant though? If a word occurs more often than others
// it is probably more indicative of the type than a word that only occurs once.
vector[hash] = 1;

So a window ["phone", "call", "phone"] produces the same vector as ["phone", "call"].

Important scope note

This is only about repetition within a single window. Repetition across documents already accumulates as counts in the per-filter-type vectors (see InMemoryVectorService / FileBasedVectorService), and that corpus-level weighting is already exercised by tests. The open question is narrowly whether a token repeated inside one short window should get extra weight.

Trade-offs

For counting within-window frequency (vector[hash] += 1):

  • A word repeated in the immediate vicinity is plausibly more indicative of the type (the TODO's intuition).
  • Moves from binary bag-of-words toward term-frequency weighting.

For keeping binary presence (status quo):

  • Windows are short (a handful of tokens per side), so within-window repeats are rare and often incidental.
  • Counting them risks letting one repeated common word dominate a single training example.
  • The ambiguous span's query vector is built by the same hash(), so switching to counts changes both the learned side and the query side; the interaction with cosine normalization is non-obvious and should be measured, not assumed.

Why low priority

The impact is small and the direction of the effect is genuinely uncertain — the change could help or hurt accuracy depending on the data. This cannot be settled by a unit test; it needs a labeled evaluation set / accuracy benchmark. Until such an evaluation exists, the status quo (binary presence) is a reasonable default.

Suggested next step

Defer until there is a disambiguation accuracy evaluation harness, then A/B the two approaches (binary presence vs. within-window term frequency) on labeled data and decide based on measured accuracy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions