Skip to content

Classify benchmark and eval failures into the four retrieval failure modes #204

Description

@SFHAJJI

What. Classify every retrieval benchmark failure and eval failure into four modes: not retrieved, wrong passage retrieved, answer buried, and right passage from the wrong document (identity failure upstream of ranking, dossier section 13). Render the counts on /benchmarks.

Why. The fourth mode is invisible to groundedness measures and is the one that produced the worst incident (section 8.3). Counting failures by mode says which brick to invest in, which is the entire point of publishing benchmarks.

Cost. A classification field on failure records, a small table on /benchmarks, goldens.

Metadata

Metadata

Assignees

No one assigned

    Labels

    axis:retrievalidentity, ranking, fusion, benchmarkenhancementNew feature or request

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions