Skip to content

Per-stage model-bindings scorecard: re-pinning a cheaper model must beat the incumbent on that stage's canaries #150

Description

@fazpu

Agreed policy (operator, 2026-07-27): one prompt-and-schema contract across models; per-model adaptation only at the transport layer (reasoning effort, provider pins, native structured-output modes — all recorded in run provenance via model_bindings); and model choice per stage decided by measurement.

The missing piece is the measurement harness: a per-stage scorecard in the eval suite so that changing any stage binding (e.g. REMEMBERSTACK_E2_EXTRACT_MODEL from gpt-5.6-luna back to a cheaper model) requires beating the incumbent on that stage's canaries — candidate coverage (#147), dated-claim rate (#146), entity/relation lint pass rate, claim quality samples.

Context: the deepseek→luna extraction comparison on LoCoMo conv-26 showed candidate recall collapse under the cheap binding (7/8 gold facts never proposed as candidates) at ~7× lower cost. Cost/quality trades per stage should be standing, gated decisions — not vibes.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GKENhTLJg1HqhbdwCmmkbc

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions