Skip to content

Repository files navigation

Metadata Matters — Quantify how catalog metadata changes text-to-SQL accuracy

CI Licence

Demo

demo

Terminal recording (docs/demo.tape via vhs): fixture dry-run and run, mm readme, then the live RQ1/RQ2 metrics. Re-record with PATH="$HOME/.local/bin:$HOME/go/bin:$PATH" vhs docs/demo.tape.

Execution accuracy by metadata level

Decoy usage by condition

Why

Text-to-SQL benchmarks score models on bare schemas, while people working from a data catalog choose tables using descriptions, sample values, business hints, and labels such as certified or deprecated. This repository measures whether that metadata changes execution accuracy and schema retrieval, including whether governance labels keep a model off plausible but wrong decoy tables. It is for catalog teams and researchers who need that comparison to be reproducible and testable offline.

Quickstart

uv run mm run rq1 -c configs/rq1.yaml --fixture

That command uses the included tiny_library fixture and the gold-SQL stub. It does not download BIRD and it does not call a hosted model. Install uv first if it is not already on the machine.

How it works

flowchart LR
  loader[BIRD layout loader] --> levels[Context M0 to M3]
  levels --> runner[Prompt, cache, and budget]
  runner --> ex[Execution accuracy]
  corpus[Column documents] --> rank[Retrieval templates]
  rank --> limited[RQ4 limited context]
  ex --> report[metrics.csv and plots]
  limited --> report
Loading

mm run loads examples through the BIRD layout, builds one context level or a retrieved column subset, asks the configured chat model for SQL, and scores execution accuracy as an order-insensitive set of rows. RQ2 scores the same SQL against a decoy copy of the database. RQ3 ranks column documents. Each run writes run.json before the metrics tables and plots.

Results

RQ Run Headline
rq1 20261002T172038Z-0d9ae03c M0 0.318
rq2 20261002T182441Z-bf771428 D0 0.351
rq3 20261002T173925Z-e0107f01 t0 sentence-transformers/all-MiniLM-L6-v2 global question 0.279
rq4 20261002T185927Z-987f2cb7 full-m2 0.512

Full tables are in each run's metrics.csv and report.md.

The live runs and the draft write-up are in docs/findings.md.

Design decisions

Roadmap

  • Confirm the BIRD licence before publishing claims that rely on that data.
  • Additional chat models or a larger sample (n: all) once budget allows.
  • Additional benchmarks, including Spider 2.0.

Licence

Apache-2.0. The BIRD dev set is not included. uv run mm download fetches it into data/bird/, which is gitignored. Confirm BIRD's licence at https://bird-bench.github.io/ before publishing results that use it. Token prices in configs/prices.yaml are maintained by hand and must be checked before a paid run.

About

Benchmark how data-catalog metadata changes text-to-SQL accuracy and schema retrieval on BIRD

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages