Skip to content

aaa-gate: whether the plane is open is a startup decision, and nothin… #39

aaa-gate: whether the plane is open is a startup decision, and nothin…

aaa-gate: whether the plane is open is a startup decision, and nothin… #39

Workflow file for this run

name: CI
on:
push:
branches: [main]
# Date tags (see CONVENTIONS §8). A tag is a claim that every lane was
# green at that commit, so tagging re-runs the full matrix rather than
# trusting that main was green when the tag was cut.
# ⚠ BOTH FORMS, because CONVENTIONS §8 and `tag.sh` mint both. A second tag
# on the same day gets `.1`, `.2`, … so names stay sortable and never
# collide — and this filter, without the second pattern, could not match
# one. `2026-08-15.1` was the first such tag ever cut here; it pushed
# cleanly and started NO run at all, which is the worst shape a tag can
# have: it asserts every lane was green while no lane ran.
#
# A trailing `*` on the first pattern would have covered it and would also
# have matched `2026-08-15-experiment`. Two explicit patterns say what is
# meant.
tags:
- "20[0-9][0-9]-[0-9][0-9]-[0-9][0-9]"
- "20[0-9][0-9]-[0-9][0-9]-[0-9][0-9].[0-9]*"
pull_request:
# Least privilege. The gate reads the repository and nothing else — it opens no
# issue, pushes no tag, publishes no package — so saying so here means a
# compromised action in the chain cannot do those things either. Borrowed from
# vercel-labs/native's Zig CI, which declares the same.
permissions:
contents: read
# One run per ref, and the superseded one is cancelled. Re-cutting a tag
# happened four times on 2026-08-15 alone, and each time the previous matrix
# kept burning four runners for work whose commit no longer existed — cancelled
# by hand, when it was noticed at all.
#
# ⚠ Deliberately keyed on the REF, not on the workflow alone: a tag run and a
# main-push run test different things and must not cancel each other.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
# Linux-only: several modules (procnet, rawsock, netlink, wireguard, …) are
# raw-syscall Linux members and are compiled by `zig build test`; socket/netns
# tests SkipZigTest cleanly on a constrained runner.
#
# ⭐ UBUNTU 26.04, PINNED, NOT `ubuntu-latest`. Two reasons, and the second is
# the one that cost real time.
#
# Pinned because `ubuntu-latest` moves under us: the image is the OTHER half of
# every live-interop test in this repository, and a gate whose oracle changes
# without a commit cannot say whether a red lane is our code or the runner's
# week. Bump this line deliberately, the way a dependency is bumped.
#
# 26.04 rather than 24.04 because of what 24.04 shipped. Its `openssh-client` is
# 9.6p1 and `mlkem768x25519-sha256` arrived in OpenSSH 9.9, so both of ssh's
# live post-quantum KEX tests could never run there — one skipped, and the other
# spawned a client that refused the algorithm, exited before connecting, and
# left our `accept` blocked until GitHub killed the job at its six-hour limit.
# That single test is what made the 2026-08-14 and 2026-08-15 matrices
# unfinishable. 26.04 ships 10.2p1, the same version as the development host,
# which both runs those tests for real and removes a whole class of
# passes-locally-fails-in-CI.
#
# ⚠ The corollary is that the runner image is now a DEPENDENCY with a version.
# When a live test starts failing, the image's release notes are the first place
# to look, not the last.
#
# TWO ARCHITECTURES, because the README makes a claim only a second one can
# check. 191 of its 225 catalog rows say Platform `any`, and a consumer reading
# that is entitled to build on aarch64. Until 2026-08-14 every lane ran on
# amd64, so `any` was a claim no job had ever tested.
#
# Three modules make it concrete rather than theoretical. `montint`, `k256` and
# `p256` gate their asm fast path on `builtin.cpu.arch == .x86_64` (see
# `asm_core.zig` / `fast_core.zig`: `pub const supported = ...`), and each
# advertises "amd64 asm + portable fallback". On an amd64 runner the dispatch
# never selects the fallback, so the path every non-amd64 consumer would run was
# the one path no lane exercised.
#
# NOT to be confused with the arm64 question closed on 2026-08-09 (`86cb149`),
# which was about PERFORMANCE -- whether `lockfree`'s seq_cst discipline costs
# anything a weaker ordering would recover. That was answered by measurement on
# TSO and needed no hardware. Portability is a different question and hardware
# is the only thing that answers it.
#
# TWO SHAPES, and the split is the point.
#
# pull_request -> `test.sh changed`, one lane. Tests the modules the
# branch touched plus their reverse-dependency
# closure, computed from `zig build module-graph`.
# push to main -> the same scoped lane, against the previous head.
# tag -> `test.sh all`, every optimize lane and both
# architectures.
#
# THE FULL MATRIX RUNS ON TAGS AND NOT ON EVERY PUSH, decided 2026-08-14 after
# the first run of it took over 95 minutes per lane. A tag is the only thing in
# this repository that CLAIMS every module passed every lane, so it is the thing
# that should pay for the claim; a push to main claims nothing of the sort and
# should not cost six runner-hours to say so. The scoped lane still gates every
# push, and it is minutes.
#
# WHY SCOPING IS SOUND, rather than merely cheaper. The driver decides the
# escalation from the GRAPH, not from which file was edited: it keeps a
# snapshot of the last known-good `module-graph` and, if any row is missing or
# altered — a module changed shape or vanished — the reverse-dep closure can no
# longer be trusted and it runs everything. With no snapshot at all it also
# runs everything. Both failure directions are fail-closed, which is what makes
# this safe to put in front of merges.
#
# The shape of the win, measured on the real graph 2026-08-12: the closure is a
# MEDIAN OF ONE MODULE, because most modules have no in-repo consumer at all. It
# reaches a few at the 90th percentile and a few dozen at its worst (`testkit`,
# which half the suite imports). A leaf change therefore costs a small multiple
# of one module's tests against the whole suite.
#
# THREE LANES, and the fourth was dropped on 2026-08-12 with a reason rather
# than a preference. `heavy_optimize` in build.zig substitutes ReleaseSafe for
# Debug on the heavy modules, which makes the old default lane's every
# (module, mode) pair a SUBSET of strict-debug's and ReleaseSafe's — it had no
# combination of its own to prove. Locally that overlap makes it nearly free to
# run after those two, since its artifacts are already built; here, where every
# lane is a separate job with its own cache, it was paying a full build for
# coverage that was already bought.
#
# The three that remain share nothing: each names one mode for every module, so
# no ordering saves a compile. They run in parallel anyway.
#
# THE LANES THEMSELVES are deliberate. Compute-heavy modules build at ReleaseSafe
# when Debug is requested (see `heavy` in build.zig — it takes `bls12_381`, the
# suite's critical path, from minutes to seconds). `-Dstrict-debug` forces it back, keeping
# CONVENTIONS §6.4's "green in Debug and ReleaseFast" honest. Integrators build
# in all three release modes; ReleaseSafe was the only one no lane covered — and
# it earned its place on 2026-08-12 by being the only lane that caught a
# use-after-scope in http's response path that every other lane passed by luck
# (see `ResponseWriter.header_buf`).
#
# They run as separate jobs so the slow one does not gate the fast one.
#
# ⚠ EVERY LANE IS CACHED, and each one measures itself. GitHub allows 10 GB of
# Actions cache PER REPOSITORY and evicts least-recently-used entries above it,
# so whether five lanes fit is an arithmetic question — and until 2026-08-15 it
# was answered with the wrong number.
#
# The retired argument ran: this repo's own `.zig-cache` is tens of GiB, so a
# lane cannot hold a useful fraction and five cannot hold one each. That
# compared a DEVELOPMENT tree against a CI job, and they are not the same
# object. The local one accumulates every optimize mode across every run ever
# made; a lane builds ONE mode, ONCE, into a fresh checkout. The real per-lane
# figure was never measured, and a cache was ruled out on an estimate that
# described something else.
#
# So the lanes cache, and each prints the size of the tree it is about to save
# (see "Report the cache footprint"). Two runs settle it: if the five together
# fit under the quota, this stays; if they evict each other, the logs will say
# so in bytes rather than in argument. A thrashing cache is worse than none —
# it costs the same wall-clock AND the pretence of being warm — so the decision
# needs a number, which is the one thing it has never had.
#
# The keys are per-lane by construction (`cachekey` in the matrix). Matrix jobs
# sharing one key clobber each other on every run; this file already records
# fixing that once, and five caching jobs is precisely when it recurs.
#
# ⚠ A red lane saves NOTHING — `actions/cache` skips its post step when the job
# fails — so the first green run is when any of this starts working, and a
# permanently red lane stays permanently cold.
#
# ⚠ AND THE PR LANE NO LONGER USES setup-zig's CACHE EITHER — it manages
# `.zig-cache` itself through `actions/cache`. Not a preference: bxp measured
# the built-in one persisting ~137 MB of a ~1.1 GB build cache with the
# compiled `o/` outputs dropped, so it restored, reported a hit, and then
# rebuilt everything. A cache that reads as warm and is cold is worse than no
# cache, because nobody goes looking for the time.
#
# That measurement was made in another repository and has NOT been repeated
# here. What is copied is the arrangement, not a verified number: if the PR
# lane's wall-clock does not fall, this is the first thing to re-measure.
#
# The old `cache-key` note applied to setup-zig's cache and is gone with it.
# The same hazard survives in a different spelling: the `key` below must stay
# distinct per caching job, since matrix jobs sharing one key clobber each
# other every run. Today only this job caches, so there is nothing to collide
# with — that stops being true the moment a second one does.
jobs:
scoped:
if: "!startsWith(github.ref, 'refs/tags/')"
runs-on: ubuntu-26.04
# A scoped lane is minutes — but this job is no longer always scoped. When
# the harness moves, `test.sh changed` escalates to the full gate here (see
# cmd_changed), and that is a run of all 225 modules whose cold-cache figure
# on a comparable lane was ~45 minutes. Thirty would have turned a correct
# escalation into a timeout, which is the worst of both: the coverage is
# paid for and then thrown away.
timeout-minutes: 75
name: changed modules + reverse-dep closure
steps:
# Full history: `test.sh changed <base>` diffs against the base branch,
# and a shallow clone has nothing to diff against.
- uses: actions/checkout@v7
with:
fetch-depth: 0
# ⚠ THE ONE ACTION STILL ON NODE 20, and not for want of bumping: v2.2.1
# is the newest `setup-zig` and its `action.yml` still says `node20`.
# `checkout` and `cache` were moved to their node24 majors on 2026-08-15;
# GitHub's deprecation warning will keep naming this one until upstream
# ships a release that does not. Nothing to do here but know why it stays.
- uses: mlugg/setup-zig@v2
with:
version: 0.16.0
# ⚠ setup-zig's OWN build cache is off on purpose, and not for the
# quota reason that keeps it off in the `full` job below. bxp measured
# it directly: it persisted ~137 MB of a ~1.1 GB build cache and
# dropped the compiled `o/` outputs — that is, everything the cache
# exists for. It reads as warm and recompiles from scratch. The Zig
# TOOLCHAIN download stays cached either way; that part works.
use-cache: false
- name: Cache Zig build artifacts
uses: actions/cache@v6
with:
path: ${{ github.workspace }}/.zig-cache
# Exact key on the sources; restore-keys falls back to the newest
# prior entry so Zig recompiles only what moved. The `base` token is
# an epoch: Zig bakes the target and CPU into its own cache key, so an
# entry built under different build flags can never hit. Bump it
# (base2, …) when this lane's flags change, to drop the dead entries.
# `base2` because the lanes moved from Ubuntu 24.04 to 26.04 on
# 2026-08-15. `runner.os` is "Linux" for both, so without a bump the
# restore-keys fallback would hand a 26.04 job a cache built against a
# different glibc and toolchain surface — a hit that is worse than a
# miss, since it is the kind that reads as warm.
key: zig-${{ runner.os }}-base2-${{ hashFiles('modules/**/*.zig', 'build.zig', 'build.zig.zon') }}
restore-keys: |
zig-${{ runner.os }}-base2-
# ⚠ THE SCOPED LANE NEEDS THE PEERS TOO, which it did not have until
# 2026-08-15. It gates every push, and until this step existed it opened
# every run with the same six-gap report the `full` job had learned to
# close — so a push touching `wireguard` or `dtls` was gated by a lane
# whose privileged and live tests skipped.
#
# It became load-bearing the moment `test.sh changed` learned to escalate
# to the FULL gate on CI when the harness moves: this job now sometimes
# runs all 225 modules, and doing that with no peers installed would
# report green over every live test in the collection.
- name: Close the environment gaps the gate reports
continue-on-error: true
run: bash scripts/ci-environment.sh
# The base differs by event, and a base that does not resolve is NOT
# allowed to read as "nothing changed" — `changed_files` fails and the
# driver escalates to the full run rather than passing having tested
# nothing. That matters here: `github.event.before` is all-zeros on a
# branch's first push.
# ⭐ THE SAME THREE FACTS THE `full` JOB PRINTS, because this lane just
# produced the question they answer. Two escalated runs of this job, same
# sources and therefore the same GitHub cache key, one restoring an older
# generation and the other hitting its primary key EXACTLY:
#
# b58ccac restore-key fallback build (all modules) 46.6 s
# dfad841 primary key, exact build (all modules) 482.9 s
#
# The exact hit was ten times slower, which cannot be explained by what
# GitHub restored — it restored more. What Zig then did with it depends on
# Zig's OWN cache key, and `native` folds a kernel range, a glibc version
# and a detected CPU into that. GitHub's amd64 pool mixes vendors: on tag
# 2026-08-15 one lane reported an AMD EPYC 7763 and two an Intel Xeon
# 8573C, in the same run.
#
# That is still a hypothesis. This job runs on every push, so putting the
# three lines here answers it in days rather than waiting for two warm tags
# to line up.
- name: What this runner is
if: always()
run: |
echo "kernel: $(uname -srm)"
cpu="$(grep -m1 -E 'model name|Model name' /proc/cpuinfo 2>/dev/null | cut -d: -f2- | xargs)"
[ -n "$cpu" ] || cpu="$(grep -m1 'CPU part' /proc/cpuinfo 2>/dev/null | cut -d: -f2- | xargs)"
echo "cpu: ${cpu:-unknown}"
echo "zig target: $(zig env 2>/dev/null | grep -m1 '\.target' | cut -d'"' -f2 || echo unknown)"
- name: Test changed modules
env:
ZIGLIBS_LANE: changed modules + reverse-dep closure
run: |
# Same as the `full` job: the imap peer is found at its default path,
# the opcua one is named by an environment variable.
export OPCUA_PYTHON="$HOME/.cache/zig-libs-opcua/bin/python3"
if [ "${{ github.event_name }}" = "pull_request" ]; then
base="origin/${{ github.base_ref }}"
else
base="${{ github.event.before }}"
fi
bash scripts/test.sh changed "$base"
full:
if: startsWith(github.ref, 'refs/tags/')
strategy:
fail-fast: false
matrix:
include:
- name: ReleaseSafe amd64
args: "-Doptimize=ReleaseSafe"
cmd: all
cachekey: relsafe
cache: true
runner: ubuntu-26.04
- name: ReleaseFast amd64
args: "-Doptimize=ReleaseFast"
cmd: all
cachekey: relfast
cache: true
runner: ubuntu-26.04
# ⚠ `build`, NOT `all` — this lane COMPILES every module in real Debug
# and runs no test, and that is the whole claim it can support. Nobody
# consumes this collection in Debug: it ships source, and the mode is
# the integrator's decision. What is worth asking of Debug is that the
# code compiles there, which is a live question, because an integrator
# developing against these modules does build them in Debug -- heavy
# ones included, and this is the only lane that ever compiles those in
# real Debug at all.
#
# Running the tests here proved nothing measurable. Debug and
# ReleaseSafe arm the same safety checks; Debug merely does not
# optimise, which makes it weaker at exposing UB, and the one
# documented cross-mode catch went the other way. Tests that run only
# in Debug: zero. Tests that SKIP in Debug: fifteen, so this lane gave
# 48/63 where its siblings gave 63/63. See `cmd_build` in
# scripts/test.sh for the measurement.
- name: strict Debug amd64 (compile only)
args: "-Dstrict-debug"
cmd: build
cachekey: strictdebug
cache: false
runner: ubuntu-26.04
# ONE arm64 lane, not three. What a second architecture buys is
# portability coverage -- alignment and endianness assumptions, atomics,
# and the asm-vs-fallback dispatch above -- and none of that varies by
# optimize mode, so the other two would mostly recompile the same
# surface. ReleaseSafe is the mode chosen because it is the one that has
# already caught a real defect here (see below). The day this lane fails
# something the amd64 ReleaseSafe lane passes, add the other two rather
# than arguing about it.
- name: ReleaseSafe arm64
args: "-Doptimize=ReleaseSafe"
cmd: all
# ⚠ NOT `relsafe-arm64`. The architecture is in the key template
# now (see the cache step), and carrying it here as well is what
# made this entry's key an extension of the amd64 lane's restore
# prefix — which is how an amd64 job came to restore this tree.
cachekey: relsafe
cache: true
runner: ubuntu-26.04-arm
name: full — ${{ matrix.name }}
runs-on: ${{ matrix.runner }}
# ⭐ A CEILING WE CHOOSE, instead of the one GitHub imposes. Twice — on
# 2026-08-14 and again on 2026-08-15 — a single wedged test held every lane
# open until the six-hour job limit killed it, burning about twenty
# runner-hours per matrix to learn nothing. The first finished matrix took
# 48 minutes wall-clock, so 90 is generous for a cold cache on either
# architecture and still fails in a tenth of the time a runaway used to.
#
# ⚠ This does NOT replace `--test-timeout` in the gate, and must not be
# thought of as doing so. A job timeout says "something here never
# finished"; the per-test deadline says WHICH test. The first is a
# backstop, the second is a diagnosis.
timeout-minutes: 90
steps:
- uses: actions/checkout@v7
# ⚠ THE ONE ACTION STILL ON NODE 20, and not for want of bumping: v2.2.1
# is the newest `setup-zig` and its `action.yml` still says `node20`.
# `checkout` and `cache` were moved to their node24 majors on 2026-08-15;
# GitHub's deprecation warning will keep naming this one until upstream
# ships a release that does not. Nothing to do here but know why it stays.
- uses: mlugg/setup-zig@v2
with:
version: 0.16.0
# setup-zig's own build cache stays off here for the same reason as
# in the scoped job: bxp measured it dropping the compiled `o/`
# outputs. The toolchain download is still cached by it.
use-cache: false
- name: Cache Zig build artifacts
# ⚠ OFF FOR THE DEBUG LANE, and the measurement is why. One cold run
# each, 2026-08-15:
#
# lane tree saved lane total
# strict Debug 4.0 GB 706 MB 271 s
# ReleaseSafe arm64 2.7 GB 687 MB 2003 s
# ReleaseSafe 2.9 GB 733 MB 2681 s
# ReleaseFast 3.1 GB 756 MB 3038 s
#
# The Debug lane holds the LARGEST tree — 225 unoptimised objects — and
# is by an order of magnitude the CHEAPEST to rebuild, because it only
# compiles and runs no test. Caching it would spend a quarter of the
# repository's quota to save about three minutes on a four-minute job,
# while the lanes that take 33 to 51 minutes compete for the remainder.
# Those three are 87-89 % compile by wall-clock, which is precisely what
# a build cache is for.
if: matrix.cache
uses: actions/cache@v6
with:
path: ${{ github.workspace }}/.zig-cache
# ⚠ `matrix.cachekey` is load-bearing. Each lane compiles EVERY module
# in ONE optimize mode, so the four trees share nothing; on a single
# key they would overwrite each other every run and every lane would
# restore someone else's mode. This file already records that bug
# once, under setup-zig's `cache-key`.
#
# ⭐⭐ AND DISTINCT NAMES ARE NOT ENOUGH — THEY MUST NOT BE PREFIXES OF
# EACH OTHER, which is a different rule and cost 42 minutes before it
# was noticed. `restore-keys` matches by PREFIX and returns the most
# recent entry that matches. The keys were `relsafe` and
# `relsafe-arm64`, so `zig-full-relsafe-` matched the ARM64 lane's
# entry, and on tag 2026-08-15 the amd64 ReleaseSafe lane restored an
# aarch64 object tree:
#
# Cache hit for restore-key: zig-full-relsafe-arm64-d9ef…
#
# Nothing was corrupt and nothing failed; the artifacts were simply for
# another architecture, so Zig rebuilt all 225 modules. That lane
# compiled for 2540 s while ReleaseFast — same runner image, same
# kernel, same CPU, same resolved target, its own uncontested key —
# compiled in 91 s. Both figures are in the same run's logs.
#
# ⚠ THIS REFUTES WHAT THIS FILE SAID ONE COMMIT AGO. The amd64 lanes
# were rebuilding despite a restore, and the explanation offered was
# that `native` bakes a kernel range, a glibc version and a CPU model
# into Zig's cache key while GitHub's amd64 pool is heterogeneous. The
# diagnostics added to prove it disproved it instead: both amd64 lanes
# report Linux 7.0.0-1011-azure, an AMD EPYC 7763, and
# x86_64-linux.7.0...7.0-gnu.2.43 — identical, and one of them was
# fast. A guess that survives because nobody measured it is worse than
# no guess, which is why that step prints three lines it does not use.
#
# The arch is therefore part of the key for EVERY lane, not a suffix on
# one of them: `relsafe-X64-` and `relsafe-ARM64-` cannot be prefixes
# of one another, whereas `relsafe-` was a prefix of `relsafe-arm64-`.
# Adding a lane means checking that its key is not a prefix of any
# other, and no longer merely that it differs.
#
# ⛔ AND WHOEVER CHANGES THE SHAPE OF A KEY MUST DELETE THE OLD
# ENTRIES. An orphaned cache does not clean itself up — it simply
# stops being reachable while still counting against the repository's
# 10 GB, so a rename pays for both generations at once.
#
# Fixing the collision above did exactly that, and by the same evening
# this repository was at 9.53 GB with 6.01 GB of it unreachable: two
# generations under `zig-full-relsafe-`, `zig-full-relfast-` and
# `zig-full-relsafe-arm64-` that no `restore-keys` prefix would ever
# look for again. With ~480 MB of headroom left, every save evicted
# something, LRU took the least recently used — which between pushes
# is the scoped lane's entry — and a lane whose key was present an
# hour earlier missed it.
#
# gh cache list # what is actually stored
# gh cache delete '<old-key>' # per entry; there is no glob
#
# Deleting the six brought it to 3.53 GB. Nothing was lost: a build
# cache regenerates, and these were already unreachable by design.
key: zig-full-${{ matrix.cachekey }}-${{ runner.arch }}-${{ hashFiles('modules/**/*.zig', 'build.zig', 'build.zig.zon') }}
restore-keys: |
zig-full-${{ matrix.cachekey }}-${{ runner.arch }}-
# scripts/test.sh all runs the format check, the hook self-test, the
# catalog/provenance gates and every module, and wraps the
# netlink-writing modules in `unshare -rn` where the runner allows it —
# otherwise those tests skip rather than fight the runner's own network
# state. It also prints a capability report naming anything the runner
# cannot do, so a lane that silently covers less than it looks like is
# visible in the log.
#
# ⚠ NO WALL-CLOCK FIGURE BELONGS IN THIS FILE. Not "none is known yet" —
# none belongs. A duration here is a number nobody re-measures and every
# reader believes: this file carried a ~1170 s figure that had never been
# rechecked, and the three that briefly replaced it were confounded by
# cache state and could not be compared with each other anyway. The run
# log is the measurement, it is per-lane and dated, and it costs nothing
# to look at. Read it there.
- name: Close the environment gaps the gate reports
# ⚠ NOT on the compile-only lane. Every one of these installs exists to
# let a LIVE test reach a real peer — a wolfSSL DTLS responder, an
# open62541 container, two Python servers — and that lane runs no test
# at all. Measured on 2026-08-15: 26 s of a 271 s lane spent installing
# things nothing would touch.
if: matrix.cmd != 'build'
# ⚠ SHARED WITH THE `scoped` JOB, deliberately. This was inline here
# until 2026-08-15, which is precisely why the lane that gates every
# push had none of it. Two copies of an install list drift; one script
# that both jobs run cannot.
continue-on-error: true
run: bash scripts/ci-environment.sh
- name: Test (${{ matrix.name }})
env:
# Names this lane's block on the run's summary page, where the four
# digests land side by side instead of in four downloadable logs.
ZIGLIBS_LANE: ${{ matrix.name }}
run: |
# The imap peer is found at its default path; the opcua one is named
# by an environment variable, so point it at the venv built above.
# A path that does not exist is not a silent loss: test.sh probes it
# and reports the gap — on the lanes that run tests. The compile-only
# lane builds no venv and runs no probe, so on that one this line is
# inert, not a safety net.
export OPCUA_PYTHON="$HOME/.cache/zig-libs-opcua/bin/python3"
bash scripts/test.sh ${{ matrix.cmd }} ${{ matrix.args }}
- name: tc action tests, as real root
# ⭐ THE LAST GAP, and the only one `unshare -rn` cannot close.
# RTM_NEWACTION checks CAP_NET_ADMIN against the INITIAL user namespace,
# so the sysctl that fixed every other netns module here does nothing
# for this one — no amount of userns permission substitutes for real
# root. scripts/test.sh therefore only PRINTS this command.
#
# That refusal is right on a development machine: `zig build` executes
# build.zig, which is arbitrary code, so a NOPASSWD rule for it would be
# passwordless root rather than a narrow grant. A hosted runner is a
# different machine — ephemeral, already executing this repository's
# code, discarded when the job ends — so the reasoning does not carry
# over, and the coverage is taken here instead of written off forever.
#
# ⚠ SEPARATE CACHE DIRECTORIES, not tidiness. Building as root into the
# job's own `.zig-cache` leaves root-owned entries that `actions/cache`
# would then save, and the next run could not write them.
#
# ⚠ This has never run anywhere. If it goes red, that is a finding about
# tc, not about this step — those tests have been skipping since the
# module was written.
#
# ⭐ `--summary all`, and it is the whole point of the step rather than
# decoration. Its first run anywhere, on tag 2026-08-15, printed NOT ONE
# LINE in thirty-five seconds and exited 0 — which is what `zig build`
# does on success, and which is indistinguishable from the tests
# skipping exactly as they always had. A step that exists to close a
# coverage gap and cannot say whether it closed it is the same fail-open
# shape it was written to fix.
if: matrix.cmd != 'build'
run: |
sudo unshare -n "$(command -v zig)" build test-tc ${{ matrix.args }} \
--summary all \
--cache-dir /tmp/zig-cache-root --global-cache-dir /tmp/zig-gcache-root
- name: Report the cache footprint and what this runner is
# The number the cache decision has never had. Whether five lanes fit
# under a 10 GB repository quota is arithmetic, and until this line
# existed the only figure anyone had was the DEVELOPMENT tree's — which
# accumulates every optimize mode across every run ever made and so
# describes a different object entirely.
#
# `if: always()` because a RED lane's footprint is worth just as much:
# actions/cache will not save it, and knowing the size anyway is how a
# permanently cold lane gets explained rather than guessed at.
#
# ⭐ AND THE KERNEL, CPU AND RESOLVED ZIG TARGET, which is how the cache
# question got answered — by refuting the explanation these lines were
# added to support.
#
# The first warm run split the lanes in two: arm64 restored, rebuilt the
# one module that had changed, and finished its compile in 42 s, while
# both amd64 lanes restored a cache of the same size and then recompiled
# essentially everything. The explanation offered was that `native` is
# not one thing — `zig env` resolves it to a kernel version range, a
# glibc version and a detected CPU, all of which feed Zig's own cache key
# — and that GitHub's amd64 pool is heterogeneous where its arm64 pool is
# not. Plausible, and wrong.
#
# These three lines said so on their first run. Both amd64 lanes reported
# Linux 7.0.0-1011-azure, an AMD EPYC 7763 and
# x86_64-linux.7.0...7.0-gnu.2.43 — identical in every respect — and one
# of them compiled in 91 s while the other took 2540 s. Nothing about the
# machine differed. What differed was which cache entry each restored,
# and that is a prefix-matching bug in the keys above, fixed there.
#
# Keep printing them. They cost nothing, they are what turned a confident
# story into a measurement, and the next surprise will want the same
# three facts about the machine that produced it.
if: always()
run: |
du -sh "$GITHUB_WORKSPACE/.zig-cache" 2>/dev/null || echo "no .zig-cache"
du -sh "$GITHUB_WORKSPACE/.zig-cache/o" 2>/dev/null || true
echo "kernel: $(uname -srm)"
# aarch64 /proc/cpuinfo has no "model name" line at all — it carries
# "CPU implementer"/"CPU part" instead — so the first run of this
# printed an empty CPU for the one lane whose hardware differs.
#
# ⚠ Tested on emptiness, NOT on the pipeline's exit status: `grep |
# cut | xargs` exits 0 even when grep matched nothing, because the
# status belongs to xargs. A `||` fallback there would never fire.
cpu="$(grep -m1 -E 'model name|Model name' /proc/cpuinfo 2>/dev/null | cut -d: -f2- | xargs)"
[ -n "$cpu" ] || cpu="$(grep -m1 'CPU part' /proc/cpuinfo 2>/dev/null | cut -d: -f2- | xargs)"
echo "cpu: ${cpu:-unknown}"
echo "zig target: $(zig env 2>/dev/null | grep -m1 '\.target' | cut -d'"' -f2 || echo unknown)"
# ⭐ THE ONE CHECK WHOSE GREEN MEANS SOMETHING.
#
# A green tick on GitHub does not say the gate ran. A SKIPPED job also shows
# as a tick, and the two shapes here exclude each other by design — `scoped`
# is skipped on tags, `full` is skipped everywhere else — so on any given ref
# roughly half the ticks are for work that never happened. Reading the run
# therefore meant knowing which jobs were supposed to run, which is exactly
# the knowledge a status check exists to spare its reader.
#
# So one job asserts the rule instead of leaving it implicit: on a tag, `full`
# passed and `scoped` did not run; on anything else, the reverse. It fails on
# `skipped` as loudly as on `failure`, because "did not run" and "ran and
# broke" are the same answer to "is this commit gated".
#
# Written down here, the rule can be argued with. Left implicit, it could only
# be misread. Borrowed from vercel-labs/native, which keeps a stable
# required-check name the same way.
#
# ⚠ THE GAP THIS USED TO CARRY IS CLOSED UPSTREAM OF HERE, and the fix is not
# in this job. A `scoped` run that detected a harness change used to print
# "this is NOT the gate", run a two-module smoke set, and exit 0 — `success`,
# and this job was happy. It happened for real on `eef1e28`, which changed
# 191 lines of opcua's driver and never built the module.
#
# `test.sh changed` now escalates to the FULL gate when it sees a harness
# change AND `GITHUB_ACTIONS` is set (see cmd_changed), so the shape this job
# cannot detect no longer occurs. That is the right place for it: this job
# reads job RESULTS, and "ran less than it should have" is not visible in a
# result. Which is worth remembering the next time something here needs
# asserting — the answer may again be in the driver rather than in this file.
gate:
name: gate
if: always()
needs: [scoped, full]
runs-on: ubuntu-26.04
timeout-minutes: 5
steps:
- name: Exactly one shape must have run, and it must have passed
env:
SCOPED: ${{ needs.scoped.result }}
FULL: ${{ needs.full.result }}
IS_TAG: ${{ startsWith(github.ref, 'refs/tags/') }}
run: |
echo "scoped=$SCOPED full=$FULL tag=$IS_TAG"
if [ "$IS_TAG" = true ]; then
ran_name=full; ran=$FULL; idle_name=scoped; idle=$SCOPED
else
ran_name=scoped; ran=$SCOPED; idle_name=full; idle=$FULL
fi
rc=0
if [ "$ran" != success ]; then
echo "::error::$ran_name is '$ran', not success — this commit is NOT gated"
rc=1
fi
# A shape that ran when it should not have means the routing above is
# wrong, and a gate whose own routing is wrong cannot vouch for
# anything it reports.
if [ "$idle" != skipped ]; then
echo "::error::$idle_name is '$idle' — it should not have run on this ref at all"
rc=1
fi
if [ "$rc" -eq 0 ]; then
echo "gate: $ran_name passed, $idle_name correctly did not run"
fi
exit "$rc"