Three places, split by what each can actually decide. Nothing blocks a push.
| where | what it checks | when | blocks? |
|---|---|---|---|
| GitHub Actions | types, the 37 suites, docs/test coverage of new API, the build, eg.py still renders |
every push and PR | the run goes red, the push already landed |
make check |
visual regression, performance baseline | when you ask for it | no |
/ship |
writes the docs, tests and examples a change owes | before committing | no |
.githooks/pre-commit |
coverage only (~12 s), opt-in | on commit, if armed | yes |
.github/workflows/ci-build.yaml already ran the suites and pyright. It now
also runs test/coverage_check.py --pushed, which reads every public symbol the
pushed range adds to the library (videocode/, src/, include/) and asks:
- is there a docstring where it is defined?
- is it mentioned in
docs/*.md? - is it mentioned in
test/?
It never writes anything, and it greps for names — so it cannot see a docstring
that lies or a test that asserts nothing. It is the floor, not the ceiling. When
it goes red, run /ship and push again. A deliberate exception is fine: say why
in the commit message.
The build job now renders eg.py too. That file is the living showcase of every
feature worth seeing, so if it stops rendering, an example in the docs has gone
stale — which is the failure mode documentation actually has.
Visual regression. The goldens in test/visual/golden/ were rendered on
Apple Silicon through Metal. The CI runner has no GPU and falls back to software
Vulkan (mesa-vulkan-drivers), whose antialiasing differs — every scene would
fail, for reasons that have nothing to do with the change. Comparing pixels
means comparing them against the same renderer.
Performance. A shared runner's timings, on software rendering, say nothing about a baseline measured on this machine. A threshold there would either never fire or fire constantly.
Both live in make check instead:
make check- Visual — renders every scene in
test/visual/scenes/and compares against its golden. Only failures not listed intest/visual/known_failures.txtcount, so "already broken" is never mistaken for "broken just now". - Performance —
test/perf/guard.pyre-runs the benchmark and compares againsttest/perf/baseline.json: load time (+30 % tolerated), render speed (+20 %), total wall time (+20 %), peak memory (+25 %). Timings keep the best of three runs; the mean drags in whatever else the laptop was doing.
Never run a blanket --update-golden to make a suite green. A golden can be
recording a bug: the matte golden did exactly that on 2026-07-28, blessing a
word whose letters were mis-spaced. Look at the image, decide whether the new
render is right, and regenerate that one scene alone.
Re-baselining performance is a decision, not a chore:
python3 test/perf/guard.py --updateDo it when a change makes the renderer legitimately slower (a feature worth its
cost) or legitimately faster (then the new floor protects the win), and say
which in the commit message. docs/optimization.md records how the current
numbers were earned — start there before accepting a regression.
CI decides whether a change is allowed in. The skill makes it worth letting
in: it reads the diff, writes the missing docstrings, the docs/*.md entries and
the tests, adds a runnable example for each, and checks whether the
surrounding documentation went stale — a wrong doc is worse than a missing one,
because it is trusted.
Deterministic checks block; judgement writes. Keeping them apart is why nothing on your critical path is slow.
If you would rather catch the coverage debt before pushing than read it off a red CI run:
git config core.hooksPath .githooks # arms .githooks/pre-commit (~12 s)
git config --unset core.hooksPath # disarms itIt runs the coverage check only — types and tests stay on GitHub, and nothing
gated to a push. git commit --no-verify bypasses it once.