Skip to content

benchmark: RS-LoCoMo-Full-v5-strong — the strong-agent protocol variant - #176

Merged
fazpu merged 1 commit into
mainfrom
feat/locomo-strong-agent
Jul 29, 2026
Merged

benchmark: RS-LoCoMo-Full-v5-strong — the strong-agent protocol variant#176
fazpu merged 1 commit into
mainfrom
feat/locomo-strong-agent

Conversation

@fazpu

@fazpu fazpu commented Jul 29, 2026

Copy link
Copy Markdown
Member

The owner-approved strong-agent variant (step 3 of the measurement plan), motivated by three controlled smoke passes on a healthy GLM-5.2 store (session recall 0.5, gold facts at rank 1) scoring 1–2/8 with the pinned gpt-4o-mini answer agent looping past the tool-call limit or returning invalid responses — the harness agent is the floor, not the memory.

A typed frozen protocol registry: full-v5 (default, bit-identical to today) and full-v5-strong (answer agent openai/gpt-5.6-luna; judge, prompts, schemas, tool catalog, budgets, temperature identical). --protocol selects at prepare; downstream stages read the pin from the run. The two fingerprints differ (answer model is in the derivation) — scores are never comparable across protocols, per the dated design note.

Review: Grok APPROVE WITH NOTES (none ship-blocking; old run dirs need re-preparing before reuse — moot, new runs prepare fresh). Final review by Claude. Fast gate local (53 passed); full matrix in CI.

@fazpu
fazpu enabled auto-merge (squash) July 29, 2026 16:22
@github-actions

Copy link
Copy Markdown

Coverage report

This PR does not seem to contain any modification to coverable code.

@fazpu
fazpu merged commit 94a28e8 into main Jul 29, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant