+| 2026-06-11 | Ranking release (develop) | develop @ 0b522a9 | -- | -- | **43.3% (13/30)**, recall@5 30/30 | **86.00% (430/500)**, recall@5 **96.60% (483/500)** | Full judged run with ship config: server `RECALL_RECENCY_BIAS=auto`, harness `temporal-answer` (`temporal_answer_hint=true`); `gpt-5-mini` answerer + pinned `gpt-5.4-mini-2026-03-17` judge; `judge_errors=0`, `memory_ingest_failures=0`. Artifact `benchmarks/results/lme_full_ship_20260611.json`, sha256 `1e53f715220d2ab4e2666106d56fd954fa4b3e4818a1bce7d060738b1bdd2d4b`. **Server-change vs harness-prompt split:** recall deltas are server-side only (prompt can't move recall); accuracy delta vs the 2026-05-17 verification run (87.00%) is within answerer noise — the two identical-config reference runs (2026-04-26 vs 2026-05-17) flip 28 answers (12 newly wrong / 16 newly right) while recall flips just 1 question, so recall is the deterministic signal and accuracy ±1pp is replicate noise. **Recall churn attribution (17 questions vs the 2026-04-26 canonical):** targeted re-runs of exactly the 17 churned questions on current main at defaults and develop at defaults show 8 of the 10 new misses also miss on current main (and the 7 newly-fixed questions already hit on current main) — i.e. 15 of the 17 churned questions (8 misses + 7 fixes) moved with this week's main merges (#191 keyword normalization), making the April canonical 97.2% floor stale; current main at defaults measures ~97.0% (485/500 est). Develop at defaults differs from current main by **1 question** (`9ea5eabc`, a near-tie rank-5/6 flip consistent with #187's deterministic timestamp tiebreak). The remaining 1 question (`00ca467f`) hits at defaults on both codebases and has no recency-trigger keyword (likely residual run noise, possibly ship-env). Attribution artifacts: `lme_churn17_main_defaults.*`, `lme_churn17_dev_defaults.*`, `analyze_churn17.py`. Failure-mode distribution vs canonical is stable (`failure_modes_ship_llm_20260612.json`: answer-construction 39 vs 41, retrieval-gap 7 vs 6); watch item: `missing-date-use` 2→6 in temporal-reasoning despite `temporal_answer_hint` (weak evidence given answerer noise). Per-category accuracy vs canonical: preference 63% vs 60%, single-session-user 94% vs 91%, assistant 100% vs 98%, multi-session 81% flat, knowledge-update 86% vs 88%, temporal 86% vs 88%. Mini-floor row (43.3% accuracy) is the judge-off develop canary, not comparable to judged rows. |
0 commit comments