Agreed policy (operator, 2026-07-27): one prompt-and-schema contract across models; per-model adaptation only at the transport layer (reasoning effort, provider pins, native structured-output modes — all recorded in run provenance via model_bindings); and model choice per stage decided by measurement.
The missing piece is the measurement harness: a per-stage scorecard in the eval suite so that changing any stage binding (e.g. REMEMBERSTACK_E2_EXTRACT_MODEL from gpt-5.6-luna back to a cheaper model) requires beating the incumbent on that stage's canaries — candidate coverage (#147), dated-claim rate (#146), entity/relation lint pass rate, claim quality samples.
Context: the deepseek→luna extraction comparison on LoCoMo conv-26 showed candidate recall collapse under the cheap binding (7/8 gold facts never proposed as candidates) at ~7× lower cost. Cost/quality trades per stage should be standing, gated decisions — not vibes.
🤖 Generated with Claude Code
https://claude.ai/code/session_01GKENhTLJg1HqhbdwCmmkbc
Agreed policy (operator, 2026-07-27): one prompt-and-schema contract across models; per-model adaptation only at the transport layer (reasoning effort, provider pins, native structured-output modes — all recorded in run provenance via
model_bindings); and model choice per stage decided by measurement.The missing piece is the measurement harness: a per-stage scorecard in the eval suite so that changing any stage binding (e.g.
REMEMBERSTACK_E2_EXTRACT_MODELfromgpt-5.6-lunaback to a cheaper model) requires beating the incumbent on that stage's canaries — candidate coverage (#147), dated-claim rate (#146), entity/relation lint pass rate, claim quality samples.Context: the deepseek→luna extraction comparison on LoCoMo conv-26 showed candidate recall collapse under the cheap binding (7/8 gold facts never proposed as candidates) at ~7× lower cost. Cost/quality trades per stage should be standing, gated decisions — not vibes.
🤖 Generated with Claude Code
https://claude.ai/code/session_01GKENhTLJg1HqhbdwCmmkbc