Conversation
Phase B of RFC 0048's rollout and the composition RFC's tie-break for spellings: measure how well a model that was never trained on GQ writes it from the schema, a one-page card and its own errors. Ground truth is computed by exact queries at a pinned graph commit, so the graph verifies itself. Records first-try validity, turns to the first correct query, task success, errors by diagnostic code, tokens and wall time per task; rows compare by value so aliases and column order do not count against a model. `run.py truth` and `report` need no model; `run` uses the Anthropic SDK with the caller's credential. Task files that name real people stay outside the repository (`tasks.example.yaml` shows the shape).
…r params and answer)
Contributor
Author
|
Closing: the instrument stays private for now. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
Adds
tools/gq-competence/: a harness that measures how well a model that was never trained on GQ writes it in context, from the schema, a one-page card and its own errors. Ground truth is computed by exact queries at a pinned graph commit, so the graph verifies itself. It records first-try validity, turns to the first correct query, task success, errors by diagnostic code, tokens and wall time; rows compare by value so aliases and column order do not count against a model.This is Phase B of RFC 0048's rollout and the instrument the composition RFC names as the tie-break for spellings (
metric(a, f)vsa.f) and the gate for new stages (#606). Draft until those RFCs are accepted or a maintainer decides tooling can land ahead of them.Backing issue / RFC
docs/rfcs/2026-09-18-gq-composition-and-language-evolution.mdand Phase B ofdocs/rfcs/0048-search-contracts.md(both drafts in rfc: search plan truth, retrieval algebra, analyzed lexical search, and the query kernel #606).Checklist
tools/, no crate changes)truthwas exercised end to end against a live graph)tools/gq-competence/README.md)Local verification
uv run tools/gq-competence/run.py truth --snapshot <commit>against a 15-task private set on the personal graph (0.11.0 server) — 15 answers computed at one pinned commitpython3 scripts/check-docs.py— OK (146 files);typos— clean;git diff --check— clean;python3 -m py_compile— OKrun.py run(the model call) — not run: no Anthropic credential on the build machineNotes for reviewers
tasks.example.yamlshows the format against the query guide's schema.claude-opus-5, adaptive thinking, efforthigh, strict tools, a cached system block carrying the card and schema.