Skip to content

tools: GQ in-context competence instrument (RFC 0048 Phase B) - #753

Closed
ragnorc wants to merge 2 commits into
mainfrom
gq-competence-instrument
Closed

ragnorc wants to merge 2 commits into
mainfrom
gq-competence-instrument

Conversation

@ragnorc

@ragnorc ragnorc commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

What & why

Adds tools/gq-competence/: a harness that measures how well a model that was never trained on GQ writes it in context, from the schema, a one-page card and its own errors. Ground truth is computed by exact queries at a pinned graph commit, so the graph verifies itself. It records first-try validity, turns to the first correct query, task success, errors by diagnostic code, tokens and wall time; rows compare by value so aliases and column order do not count against a model.

This is Phase B of RFC 0048's rollout and the instrument the composition RFC names as the tie-break for spellings (metric(a, f) vs a.f) and the gate for new stages (#606). Draft until those RFCs are accepted or a maintainer decides tooling can land ahead of them.

Backing issue / RFC

Checklist

  • Change is focused (one tool under tools/, no crate changes)
  • Tests added/updated for behavior changes (N/A: measurement tooling; truth was exercised end to end against a live graph)
  • Public docs updated if user-facing surface changed (tools/gq-competence/README.md)
  • Reviewed against docs/dev/invariants.md — no engine change

Local verification

  • uv run tools/gq-competence/run.py truth --snapshot <commit> against a 15-task private set on the personal graph (0.11.0 server) — 15 answers computed at one pinned commit
  • python3 scripts/check-docs.py — OK (146 files); typos — clean; git diff --check — clean; python3 -m py_compile — OK
  • run.py run (the model call) — not run: no Anthropic credential on the build machine

Notes for reviewers

  • The private task set for the personal graph is not in the tree (it names real people); tasks.example.yaml shows the format against the query guide's schema.
  • The card describes today's grammar. Comparing a proposed spelling means a second card and the same tasks, model, schema and snapshot; the README says how.
  • Model defaults follow the SDK guidance: claude-opus-5, adaptive thinking, effort high, strict tools, a cached system block carrying the card and schema.

Phase B of RFC 0048's rollout and the composition RFC's tie-break for
spellings: measure how well a model that was never trained on GQ writes it
from the schema, a one-page card and its own errors. Ground truth is
computed by exact queries at a pinned graph commit, so the graph verifies
itself. Records first-try validity, turns to the first correct query, task
success, errors by diagnostic code, tokens and wall time per task; rows
compare by value so aliases and column order do not count against a model.

`run.py truth` and `report` need no model; `run` uses the Anthropic SDK with
the caller's credential. Task files that name real people stay outside the
repository (`tasks.example.yaml` shows the shape).
@ragnorc

ragnorc commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Closing: the instrument stays private for now.

@ragnorc ragnorc closed this Sep 18, 2026
@ragnorc
ragnorc deleted the gq-competence-instrument branch September 18, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant