production foundations: a gate that can fail, an exact environment, and an API that cannot quietly do the wrong thing #46
Workflow file for this run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # The fidelity gate, run on a pinned Linux oracle. | |
| # | |
| # Determinism has two halves and this workflow needed both: | |
| # | |
| # inputs 16 PDFs frozen in testkit/fixtures/, pinned by SHA-256. They | |
| # used to be regenerated here, so a Chromium update on the runner | |
| # silently changed a corpus document and moved a gated metric 5x. | |
| # environment the renderer decides the numbers. `evidence.environment()` | |
| # fingerprints OS, Python minor, LibreOffice, the metric fonts and | |
| # every measurement dependency, and `canonical` means that exact | |
| # combination -- it used to mean `os == "linux"`, which cannot | |
| # tell Chromium 149 from 150 or Python 3.12.3 from 3.12.13. | |
| # | |
| # Why Linux is canonical: the fidelity numbers depend on the renderer's fonts. | |
| # Local Windows runs render with real Arial/Times New Roman; this runner renders | |
| # with Liberation metrics-compatible substitutes. Same code, different wraps. One | |
| # environment has to be the reference, and it must be the one everyone can | |
| # reproduce -- so CI is the number of record and local runs are indicative. | |
| # | |
| # Two lanes (testkit/runall.py): `raw` is the uncontaminated converter number, | |
| # `product` is exactdoc.options.PRODUCT -- the profile the API, the CLI and every | |
| # published number all share. refine() tunes against the same renderer the gate | |
| # measures with, so only the pair is meaningful, and **both** lanes gate the exit | |
| # code. Gating on the refined lane alone meant the control lane, whose entire | |
| # purpose is to be untainted, was the one nobody had to answer for. | |
| # | |
| # Every step here is fail-closed. That is the whole design: a green check must | |
| # mean that all 16 manifest documents existed, the renderer answered, every | |
| # required metric was computed, nothing regressed past its recorded number, and | |
| # the backend policy still describes reality. It previously could mean none of | |
| # those things -- see the docstrings in testkit/gate.py for the list, each entry | |
| # of which is now a test in tests/test_gate_mutations.py. | |
| # | |
| # Provisioning is scripts/bootstrap.sh, the same command a contributor runs, so | |
| # CI cannot drift away from the documented setup without going red. --strict | |
| # makes a missing oracle a failure. | |
| # | |
| # The dependency versions come from uv.lock (--frozen). The goldens are pinned | |
| # to the PyMuPDF version -- measured: 1.26 and 1.24 both put 02_research_paper | |
| # p2 at 4 blocks where 1.28 puts 7 -- so an unpinned resolve would fail the | |
| # golden step for a reason that has nothing to do with this repository's code. | |
| name: gate | |
| on: | |
| push: | |
| branches: [main] | |
| pull_request: | |
| jobs: | |
| gate: | |
| runs-on: ubuntu-24.04 | |
| timeout-minutes: 90 | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - uses: actions/setup-python@v5 | |
| with: | |
| python-version: "3.12" | |
| - name: Install uv (uv.lock is the pinned truth) | |
| uses: astral-sh/setup-uv@v5 | |
| with: | |
| enable-cache: true | |
| - name: Provision the oracles (LibreOffice, Chromium, fonts) + deps | |
| run: bash scripts/bootstrap.sh --strict | |
| # Restrict the renderer to the pinned font set for every later step. A | |
| # runner image ships a large font collection; with the corpus already | |
| # frozen byte-for-byte, that alone still moved c4_i18n's dy_p50 from | |
| # 0.15pt to 2.1pt, because LibreOffice resolved its CJK and RTL runs to | |
| # faces the measurement environment does not have. See scripts/fonts.conf. | |
| - name: Pin the font environment | |
| run: | | |
| mkdir -p /tmp/exactdoc-fontconfig | |
| echo "FONTCONFIG_FILE=$GITHUB_WORKSPACE/scripts/fonts.conf" >> "$GITHUB_ENV" | |
| FONTCONFIG_FILE="$GITHUB_WORKSPACE/scripts/fonts.conf" fc-list : family \ | |
| | tr ',' '\n' | sort -u | |
| # The metric corpus is 16 PDFs frozen in testkit/fixtures/ and pinned by | |
| # SHA-256. It is NOT regenerated here, and that is the fix for the failure | |
| # this workflow actually had: the baseline was recorded against a corpus | |
| # built with Chromium 149, this runner ships Chromium 150, `c4_i18n` came | |
| # out a different document, and its vertical drift moved 0.15pt -> 0.7pt. | |
| # A gated metric moved 5x because of a browser update. A generated corpus | |
| # cannot be a measurement baseline. | |
| - name: Corpus fixtures - byte-identical to the record? | |
| run: uv run python testkit/corpus_manifest.py verify | |
| # The generators still run, and no longer gate a measured number. Drift | |
| # between a fresh generation and the frozen bytes is reported: it says the | |
| # toolchain moved, which is worth knowing and is not a regression here. | |
| - name: Corpus generators still work (drift reported, not gated) | |
| run: uv run python tests/test_corpus_generation.py | |
| - name: Unit tests (write purity, corpus degradation, gate mutations) | |
| run: | | |
| uv run python tests/test_purity.py | |
| uv run python tests/test_corpus_degradation.py | |
| uv run python tests/test_gate_mutations.py | |
| # The permissive runtime boundary, and the reason the licence flip is a | |
| # real change rather than a metadata edit. This makes `fitz` unimportable | |
| # and then converts the representative fixtures, which is stricter than a | |
| # virtualenv without the package: it also catches an import that something | |
| # else in the interpreter has already performed. | |
| - name: Convert with PyMuPDF made unimportable | |
| run: uv run python tests/test_no_pymupdf.py | |
| - name: Golden IR - the parser's output must not drift | |
| run: uv run python testkit/golden_ir.py verify | |
| - name: Fidelity gate, both lanes, fail closed | |
| run: uv run python testkit/runall.py | |
| # No longer continue-on-error, and it is EXPECTED TO FAIL right now. That | |
| # combination is deliberate. | |
| # | |
| # It was reporting-only "until the swap lands", which made the number it | |
| # exists to drive the one number nothing depended on. The policy now lives | |
| # in testkit/parity_policy.json with numeric floors, so the executable rule | |
| # and the ratified rule are the same rule. | |
| # | |
| # Applying that rule to every dimension independently -- including vertical | |
| # drift, which the old comparison never looked at -- found 2 unwaived | |
| # regressions (05_memo and f1_fpdf_brief, both dy_p50) that had been | |
| # reported as "same". They are attributed to the same core-14 font-metric | |
| # cause as the four D2 waivers and are NOT waived, because widening a | |
| # waiver from four documents to six is a product decision. | |
| # | |
| # Do not re-add continue-on-error to make this green. Red is the correct | |
| # state until the shortfalls are ratified or fixed; going green by ignoring | |
| # the result is precisely how this step stopped working the first time. | |
| - name: Backend parity - the licence-swap verdict | |
| run: uv run python testkit/backend_parity.py | |
| - name: Evidence - one artifact every published number traces to | |
| if: always() | |
| run: uv run python testkit/evidence.py --out testkit/batch/evidence.json | |
| - name: Upload lane results and evidence | |
| if: always() | |
| uses: actions/upload-artifact@v4 | |
| with: | |
| name: gate-results | |
| path: | | |
| testkit/batch/evidence.json | |
| testkit/batch/lane_raw/results.json | |
| testkit/batch/lane_raw/verdict.json | |
| testkit/batch/lane_product/results.json | |
| testkit/batch/lane_product/verdict.json | |
| if-no-files-found: error |