What. Classify every retrieval benchmark failure and eval failure into four modes: not retrieved, wrong passage retrieved, answer buried, and right passage from the wrong document (identity failure upstream of ranking, dossier section 13). Render the counts on /benchmarks.
Why. The fourth mode is invisible to groundedness measures and is the one that produced the worst incident (section 8.3). Counting failures by mode says which brick to invest in, which is the entire point of publishing benchmarks.
Cost. A classification field on failure records, a small table on /benchmarks, goldens.
What. Classify every retrieval benchmark failure and eval failure into four modes: not retrieved, wrong passage retrieved, answer buried, and right passage from the wrong document (identity failure upstream of ranking, dossier section 13). Render the counts on /benchmarks.
Why. The fourth mode is invisible to groundedness measures and is the one that produced the worst incident (section 8.3). Counting failures by mode says which brick to invest in, which is the entire point of publishing benchmarks.
Cost. A classification field on failure records, a small table on /benchmarks, goldens.