Skip to content

Commit 41bf8d0

Browse files
jack-arturoclaude
andcommitted
docs(bench): log full judged 500q LongMemEval ship-config run with churn attribution
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1 parent 0b522a9 commit 41bf8d0

1 file changed

Lines changed: 1 addition & 0 deletions

File tree

‎benchmarks/EXPERIMENT_LOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -60,6 +60,7 @@ Detailed experiment history:
6060
| 2026-05-17 | Publication verification | feat/automem-arxiv-publication | **85.20% (259/304)** | **84.74% (1683/1986)** | **70.00% (21/30)** | **87.00% (435/500)** | Fresh local publication reruns. LoCoMo full used pinned `gpt-5.4-mini-2026-03-17` judge, 444 judge calls, 0 skips/errors, estimated judge cost `$0.7909`, artifact `benchmarks/results/locomo_baseline_20260517_193934.json`, sha256 `a75816e9a6d3302c22b34852b75ac19a9d9f5cb27d1a109e0af7e49359330716`. LongMemEval full used `gpt-5-mini` answerer + `gpt-5.4-mini-2026-03-17` judge, recall@5 **97.00% (485/500)**, `memory_ingest_failures=0`, `judge_errors=0`, `publishable=true`, artifact `benchmarks/results/longmemeval-full-publication-20260518.json`, sha256 `ed6f7cf69b7be6fa0050536ec2b0f947f5510afd8c2a374b3fafb9cde009da75`. |
6161
| 2026-06-06 | main-refresh (no judge) | main @ b1df86c | **83.40% (196/235)** | -- | -- | -- | Same local `.env`, snapshot eval after #173. Comparison anchor for PR #124/#72 hardening; cat-5 judge disabled, so 69 complex questions are skipped and the result is directional. |
6262
| 2026-06-06 | PR #124 + #72 hardening | feat/entity-identity-hardening | **83.40% (+0.0)** | **81.71% (1260/1542)** | -- | -- | Entity quality gates, safe Entity-node migration/dedup, current-state identity synthesis, and disabled scheduled synthesis. Flat vs same-env main on LoCoMo-mini; full run is judge-off with 444 cat-5 questions skipped, so focused graph/entity regressions carry the entity-pollution risk. |
63+
| 2026-06-11 | Ranking release (develop) | develop @ 0b522a9 | -- | -- | **43.3% (13/30)**, recall@5 30/30 | **86.00% (430/500)**, recall@5 **96.60% (483/500)** | Full judged run with ship config: server `RECALL_RECENCY_BIAS=auto`, harness `temporal-answer` (`temporal_answer_hint=true`); `gpt-5-mini` answerer + pinned `gpt-5.4-mini-2026-03-17` judge; `judge_errors=0`, `memory_ingest_failures=0`. Artifact `benchmarks/results/lme_full_ship_20260611.json`, sha256 `1e53f715220d2ab4e2666106d56fd954fa4b3e4818a1bce7d060738b1bdd2d4b`. **Server-change vs harness-prompt split:** recall deltas are server-side only (prompt can't move recall); accuracy delta vs the 2026-05-17 verification run (87.00%) is within answerer noise — the two identical-config reference runs (2026-04-26 vs 2026-05-17) flip 28 answers (12 newly wrong / 16 newly right) while recall flips just 1 question, so recall is the deterministic signal and accuracy ±1pp is replicate noise. **Recall churn attribution (17 questions vs the 2026-04-26 canonical):** targeted re-runs of exactly the 17 churned questions on current main at defaults and develop at defaults show 8 of the 10 new misses also miss on current main (and the 7 newly-fixed questions already hit on current main) — i.e. 15 of the 17 churned questions (8 misses + 7 fixes) moved with this week's main merges (#191 keyword normalization), making the April canonical 97.2% floor stale; current main at defaults measures ~97.0% (485/500 est). Develop at defaults differs from current main by **1 question** (`9ea5eabc`, a near-tie rank-5/6 flip consistent with #187's deterministic timestamp tiebreak). The remaining 1 question (`00ca467f`) hits at defaults on both codebases and has no recency-trigger keyword (likely residual run noise, possibly ship-env). Attribution artifacts: `lme_churn17_main_defaults.*`, `lme_churn17_dev_defaults.*`, `analyze_churn17.py`. Failure-mode distribution vs canonical is stable (`failure_modes_ship_llm_20260612.json`: answer-construction 39 vs 41, retrieval-gap 7 vs 6); watch item: `missing-date-use` 2→6 in temporal-reasoning despite `temporal_answer_hint` (weak evidence given answerer noise). Per-category accuracy vs canonical: preference 63% vs 60%, single-session-user 94% vs 91%, assistant 100% vs 98%, multi-session 81% flat, knowledge-update 86% vs 88%, temporal 86% vs 88%. Mini-floor row (43.3% accuracy) is the judge-off develop canary, not comparable to judged rows. |
6364

6465
### Category Breakdown (LoCoMo-mini)
6566

0 commit comments

Comments
 (0)