Skip to content

Latest commit

 

History

History
1004 lines (610 loc) · 227 KB

File metadata and controls

1004 lines (610 loc) · 227 KB

CLAUDE.md

Project conventions for Claude (Anthropic CLI / agent sessions) working on this repository. Sessions inherit nothing across cold starts, so the durable rules live here.

A recurring value across these conventions: prefer the existing structure over indirection that narrows nothing — churn (renames, façades, per-module "engines", parallel counters/diagnostics) must earn its keep. Where that trade-off bites, the relevant section calls it out as the bad-façade / namespace-churn trap.

Priority when anti-churn and consistency collide: consistency wins when the inconsistency is recurring and reader-facing (every new reader re-pays it); anti-churn wins only when the change is purely cosmetic — a rename or reshuffle that removes no confusion. "Earn its keep" is not counted in functional gain alone: a persistent inconsistency is a real, compounding cost, not a cosmetic one. This leans consistency-first on purpose, because the project is now mature — low debt, high coverage, the #563-era refactor closed — and churn is cheap on clean, well-tested code, so a consistency fix that disturbs the module-boundary baseline or a namespace is usually worth it. (This is a shift from the refactor era, when churn fought an in-flight baseline and stability rightly came first.)

A name that encodes a VALUE or a SCOPE re-breaks every time the value or the scope moves, and nothing reports it (#1525). This is the predictive half of the rule above: that one says when a standing inconsistency is worth fixing, this one says which names will become one. Measured on this repository's own history, in three shapes:

  • Named by its value. The margin utilities were .ffc-mt-15 / .ffc-ml-18, so #1176's one-step change made nine class names lie at once; they read the step now (.ffc-mt-xl). Two tokens went further and were deleted rather than renamed — --ffc-spacing-5 and --ffc-spacing-15 were not steps but what remained of a second ladder, and the naming rule is what made that debt leave instead of hiding behind a fresh name.
  • Named by its scope. A sheet named for one screen that loads on every FFC screen, and one file living under two handles — ffc-admin-submissions.css and ffc-admin.css, both described where they bite, under "Stylesheet architecture" and "Page scope on the admin wrap". Cited here rather than re-explained, because the mechanism is already written down twice and a third copy is what this file's own redundancy looks like.
  • Named after a neighbour. The table is ffc_short_urls; ffc_url_shortener is only an option-key and metabox-id prefix. Neither name is wrong on its own, which is exactly why the pair costs a lookup every time.

The test is one question, asked before the name is chosen: what would have to change for this name to become false? When the answer is a number, a screen list, or a sibling's spelling, the name is borrowing something it does not own. Prefer the step over the pixel, the family over the screen, and the thing's own identity over its neighbour's.

Two limits, both hard. A name that is a stored value does not rename — sindicato, rf_encrypted and every field_key live in rows, and the criterion under "Domain conventions" governs. And a rename is the most expensive change here to get wrong (the module-boundary baseline, @covers, alias mocks, the autoloader), so it rides a PR that already touches the file; a rename-only sweep is the churn the rule above refuses.

The mechanical half of this is already clean and needs no rule: a method named for a read that writes. Measured across includes/, three — two getters memoising through a transient and one flash message that is read-once-and-clear by definition. Nothing to fix, so nothing is written down about it; a rule for a defect class that does not exist is the speculation this file refuses everywhere else.

A number in this file is a claim about a value something else owns, and it goes stale in silence (#1261). Nothing warns anybody: the PR that shrinks a ratchet edits the constant, never the paragraph in another section that describes it — the guard still passes, because a docblock is not an assertion. Measured on this file, the class is live in three shapes at once. A contradiction: "Quality gates and testing" said the CssNamespaceAnchorTest baseline was empty (true) while "Stylesheets and theme" said "the three survivors … stay on purpose" (false — #1202 renamed them, which is what emptied it). A stale count: ffc-common.css was described as 1,226 lines carrying 135 tokens across 206 selectors; measured, 1,361 / 91 / 232. And the worst shape, an unverifiable one: neither 135 nor the theme section's 78 matches any reading of the file anybody can reproduce, so the method was never written down and the number could not be checked even when it was written.

So: state the invariant, not the reading. "The baseline blocks at zero" survives a shrink; "the baseline holds 3 entries" does not. "The dark block redefines the colour tokens and only the colour tokens" survives a new token; "57 of the 78" does not. Where the count genuinely carries the lesson — the 60 → 51 → 3 → 0 arc is the story — keep it and write it as history, in the past tense, which cannot go stale; the contradiction above was a history sentence written in the present. And where a number stays, make it re-measurable: say what was counted (91 tokens declared on :root), because a figure whose method is unrecorded cannot be confirmed or refuted by the next reader. A count that drifts with every edit and carries nothing — a line count, a selector count — is better deleted than refreshed.

All 517 of them were enumerated once (#1261), and the ratio is the part worth knowing: most numbers here are right. The sweep read every numeric token in the file, dropped the ones that claim nothing measurable (list markers, section numbers in the table of contents, an example IP, an HTTP status), and traced each survivor to the constant or file that owns it. Roughly forty were confirmed against their owner and left alone — the palette's :root tokens and colour values, the count the dark block redefines, the 13 extends AbstractRepository classes split 4 + 3 + 6, the screen map, MIN_REASON_LENGTH, the PHP and WordPress floors, the MariaDB pin, the seven typography steps and the nine spacing rungs, the four vendored bundle versions against their filenames, the 16 methods of BatchedExportSourceInterface, every ratchet register that blocks at zero. So the rule is not "distrust every number" — it is that the ones which go stale do so in a predictable place: a count of files, of classes, of findings a linter reports, of anything a later PR edits without reading this file. Four of those confirmed figures had drifted again within four releases (#1525) — the palette's three counts and the screen map, each because a later PR added a token or a screen without reading this file, and the @covers file count beside them. Every one survives here with its number removed rather than refreshed, which is what the paragraph above prescribes: the palette's invariant was re-measured and holds exactly, so the counts were carrying nothing it did not already say.

A stale number is a fact; a stale number in a sentence that also states a conclusion is a lie. Three of that sweep's corrections were of the second kind, and each told a reader something untrue about the project's state rather than merely something out of date. ficha was described as an open product decision the day after #1280 decided it — the drift window is one merged PR, not a season. Five bare element selectors were described as "already debt in #1152's baseline" when the baseline had been empty since #1202 and the five had been anchored, not merely re-counted. A table row called the utility layer "drifting" and listed nine components and three dead classes that #1171 had already moved and deleted. In each case the number was the symptom and the claim was the damage — so when a figure looks stale, read the whole sentence before refreshing it, because the fix is often to rewrite the claim and delete the figure.

A section number is a number too, and it lives in other files. The renumbering that gave the CSS arc its own section (#1288, three PRs before this one) broke six CLAUDE.md §N references inside the codebase — three pointing at the legacy-shim rules, three at the domain conventions — and nothing reported it, because a docblock is not an assertion and no gate reads this file. The fix is not to renumber them: it is that a cross-reference names the section, never numbers it (CLAUDE.md "Security & PII conventions"), which survives any reordering. All nine references in the tree were converted, including the three that happened to still be correct, because being correct today is what a latent one always looks like. The conversion stopped at the tree's edge, though, and this file kept twelve §N references of its own — five with a section name beside the number, which survives a renumber carrying a wrong number, and seven bare. Converted in the same pass that found the 101 above, and for the same reason: a rule a document does not follow is a rule a reader learns to discount.

Language: every versioned artifact is in English; the conversation is not. Code, comments, docblocks, assertion messages, CHANGELOG.md, commit messages, issue titles and bodies, and PR titles and bodies — English. The i18n source strings too, with the translation living in languages/ — that is what keeps the plugin translatable into any language and not just the one its maintainer speaks. The conversation with the maintainer is primarily in Brazilian Portuguese, and that does not decide the language of what stays in the repository — writing this down is the fix for how it drifted: nothing here said so, so three releases' worth of CHANGELOG, 24 test files and 75 assertion messages came out in Portuguese without anyone deciding to (#1260). Writing it down was necessary and not sufficient — the cleanup it triggered then shipped incomplete three times, so the comment and source-string halves are now enforced by CommentLanguageTest ("Quality gates and testing" → "Comment-language guard"); what that guard deliberately cannot reach is named there.

Issues and PRs are the half that is easiest to forget, and the drift there is not historical — it happened while this very rule was being written. Every PR body and every issue opened in the session that introduced this paragraph was in Portuguese, including the PR that introduced the paragraph and the issue tracking the rest of the cleanup. The reason is mechanical and will repeat: a PR body is composed in the flow of the conversation, so it inherits the conversation's language unless something interrupts, while a code comment is composed in the file and inherits the file's. So the discipline belongs at the moment of writing the body, not at review — nothing else catches it, and by review the text is already the thing being approved. What already merged stays as it is: a merged PR body is a record of a moment, and editing dozens of them would make each PR title disagree with its own squash subject on main, which cannot be rewritten (#1260).

That last half is the one to measure carefully, because a narrow scan will tell you it is clean when it is not. A first pass over the 7,281 source strings reported zero in Portuguese; it was looking for a short list of function words, so it missed a whole Portuguese sentence shipped as a source string ('Nao foi possivel decifrar %1$s do usuario %2$d…', in the key-rotation migration — fixed in #1260) and could not have seen the systemic case below. Widen the dictionary, then filter by what a Portuguese word actually indicates: a domain acronym inside an English sentence ("Search by name, email, ID or CPF/RF") is correct — CPF is a proper noun here, like VAT — while a functional word (não, são, pelo, uma) only appears in a Portuguese sentence. The difference between the two filters is 216 hits against 61.

ficha was the open case, and it closed the way this file says such cases close — by a product decision, not a mechanical rename. It was an inconsistency rather than a translation defect: about 55 source strings used it inside otherwise-English text ("Download Ficha"), while the codebase already said three different things at once — the generator's docblock read "ficha (data sheet)", ReregistrationAdmin mapped 'ficha' => __( 'Record' ), and FilenameHelper emitted _x( 'ficha', 'pdf filename prefix' ), a Portuguese source string whose translation should have gone the other way. #1280 settled it on Record (#1264). What survives is the criterion below, and exactly two source strings that still say "ficha" on purpose: one names the template kind, which is a stored value, and one records the hook family the rename came from. Neither is debt — grep for a third before assuming otherwise.

The exception is a criterion, not a list: a domain term stays in Portuguese when renaming it would break the database or require a migration. rf_encrypted / rf_hash / rf_normalized are real columns in a CREATE TABLE; sindicato, acumulo_cargos, endereco_numero, email_institucional, termo_ciencia and nome_completo are field_key values written into ffc_custom_fields and keys of the data JSON of every stored submission. Renaming those means migrating every install's rows, for a term like RF (Registro Funcional) that has no English equivalent an SME operator would recognise. Everything else renames — a method name, a class, a local variable touch no stored row, so get_sindicato_options() is not covered by the exception even though the field key it serves is. Expect that split to look odd in the diff (get_union_options() feeding 'field_key' => 'sindicato') and to need one line saying the key is the stored value; that is the honest shape, not an oversight.

Table of contents

  1. Contributing workflow — git / PR / release: pull-request workflow, branch naming, develop-branch workflow, versioning, CHANGELOG conventions, what not to do.
  2. Quality gates and testing — the CI gate list + coverage floors, the guards (one subsection each), config-level exclusions, test infrastructure, build & assets.
  3. Architecture and patterns — repository pattern, module bootstrap (loaders), shared-service directories, email pipeline, CSV export, captcha, batched migrations.
  4. Stylesheets and theme — how many sheets and why, naming and composition, page scope on the admin wrap, the light/dark palette.
  5. Domain conventions — date/time storage, settings reads and writes, admin number inputs, capability naming, security & PII.
  6. Legacy and tech debt — compat shims + evidence-gating, deprecation cycles and the @removal marker.

1. Contributing workflow

Pull-request workflow

The repository has "Allow auto-merge" enabled in Settings → General. Use it on every PR — no manual squash + merge unless auto-merge fails.

PR base branch: by default, target develop, not main. The only PRs that target main are (1) the periodic release PR develop → main that consolidates the accumulated batch with a single version bump, and (2) hotfix PRs from hotfix/* branches when a critical bug needs to ship without waiting for develop's queue to consolidate. See "Develop branch workflow" below for the full mapping.

After mcp__github__create_pull_request:

  1. mcp__github__update_pull_request with draft: false (auto-merge only fires on non-draft PRs).
  2. mcp__github__enable_pr_auto_merge with mergeMethod: SQUASH.
  3. End your turn. GitHub merges as soon as every required check passes; <github-webhook-activity> fires the merge event back to the session.

Don't poll CI manually after step 2 unless the user asks. The webhook subscription delivers failures + comments + the final merge event; green completion is silent by design.

Batch related work; gate the draft→ready flip. Group trivially-related changes — especially docs-only (CHANGELOG / CLAUDE.md / comments) — into one PR; they gain nothing from the per-PR testes deploy and only multiply CI runs and rebase churn. Use draft to accumulate commits and keep the window short: open the PR late (work essentially done + locally green), rebase if develop moves under it, and treat a second forced rebase as the signal to finish or split. Before flipping draft→ready + auto-merge: if the PR completes a pre-agreed unit (a planned sprint/roadmap item with obvious done-criteria), go straight to ready + auto-merge; if the scope is ad-hoc or emerged mid-session — completeness not obvious from a plan — confirm with the user first, so follow-ups stay batched in the same PR instead of spawning a trickle of tiny PRs. No calendar deadline on drafts — the signal is drift (develop moving under the branch), not elapsed time.

The cost of ignoring that first sentence, measured (#1260). A comment-only translation of 27 test files was split into four PRs because 56 KB of prose is not reviewable in one diff. The size argument holds; four did not. Two of them hit the same CHANGELOG.md conflict and each cost a rebase, which is precisely the "rebase churn" the rule names. The reviewable unit is the diff, so split by size when a single diff would genuinely be unreadable — and then stop, rather than splitting again for tidiness. Where the work is one mechanical pass over many files, two large PRs beat four comfortable ones.

Closing an issue — leave no false-positive checkboxes. When an issue closes as completed, its acceptance-criteria / delivery checkboxes must reflect reality: either tick every delivered box (- [x]) or leave a short closing comment that maps each delivered item to the PR that shipped it. An issue that closes with unticked - [ ] boxes for work that actually shipped reads as pending to a later scan — this is exactly the false positive reconciled across #249, #647, #649, #650, #697 (delivered work split across separate PRs, boxes never revisited). Legitimately-deferred boxes stay unticked, but the closing comment must name them as deferred (see #711's engine deferral) so "unticked" is never ambiguous between done and dropped.

Closes #NNN does not auto-close in this workflow — close feature issues by hand. GitHub only auto-closes a linked issue when the PR merges into the repository's default branch, which is deliberately kept as main (see "Develop branch workflow"), not develop. So a feature PR that merges to develop with Closes #NNN in its body leaves the issue open. Close it manually — with the checkbox reconciliation above — as soon as its work lands on develop (matching the repo's precedent that issues close on work-complete, e.g. #711 / #591 / #249, not on release-to-prod); do not wait for the release PR to carry it to main, or done work lingers as an open, unticked issue (the #728 case). Genuinely-future work (a removal scheduled releases out, like #730) correctly stays open — this rule is about work that has actually shipped to develop.

Branch naming

Use claude/<short-kebab-description> for feature branches that target develop. Examples in main history: claude/js-coverage-sprint-B-audience-smoke, claude/csv-import-normalize-cpf-rf. Use hotfix/<short-kebab-description> when cutting an urgent fix from main (see "Develop branch workflow" → Hotfixes). Always go through a PR — the PR is what gets the commit through CI, which a direct push does not.

What an agent's credential can actually do, measured on the 6.25.0 release — because this was written down wrong for a long time. This section used to claim the remote refuses pushes to main and develop. It does not: the post-release sync was a git push --force-with-lease origin develop by an agent and it went through, reporting Bypassed rule violations for refs/heads/develop: 12 of 12 required status checks are expected. The ruleset exists; the credential has bypass. So "go through a PR" is a rule about what you should do, never a wall that stops you — and a false wall is worse than no wall, because it stops people from checking.

The real limit is a different one, and it is not written anywhere else: an agent cannot delete a ref at all. git push --delete returns HTTP 403 on any branch, feature branches included. Confirmed by probe — a throwaway branch was created solely to separate "this branch is protected" from "no deletion passes", and it is the second. So the credential creates and updates refs, including by force, and never deletes. Deleting a merged feature branch is the maintainer's to do; an agent that promises otherwise cannot deliver, and leaves the branch behind.

Develop branch workflow

Adopted v6.7.7 to decouple iteration cadence from production. Source domain (prod) only sees one version bump per batch; the testes domain runs develop HEAD and absorbs the per-PR churn.

Branch topology

main      ─o──────────────────────o────────────  ← PROD (source domain)
                                  ↑
                                  release PR
                                  with consolidated bump
develop   ─o─o─o─o─o─o─o─o─o─o─o─/             ← TESTES (testes domain)
           PR1 PR2 PR3 …                         each merge auto-deploys
  • main — the production branch. Updated only by (1) the release PR develop → main (squash merge with the final bump + consolidated CHANGELOG entry) and (2) hotfix PRs hotfix/* → main when a critical bug bypasses the queue.
  • develop — the integration branch. Default base for every feature PR. Each merge into develop triggers .github/workflows/deploy-develop.yml, which rsyncs the working tree to the testes server.
  • hotfix/* — short-lived. Cut from main, merge back to main, then rebase develop on top of the new main (see Hotfixes).

Per-PR flow (the common case)

  1. Feature branch claude/<desc> cut from develop.
  2. PR targets develop. CI gates run identically to PRs against main.
  3. No FFC_VERSION bump in the feature PR. Develop accumulates work under the existing FFC_VERSION until release time.
  4. CHANGELOG entries go under the [Unreleased] section at the top of CHANGELOG.md — they stay there across multiple PRs.
  5. Auto-merge enabled (SQUASH). Once merged, deploy-develop.yml pushes the new HEAD to the testes server within ~1 minute.

Dependabot PRs

Dependabot is configured with target-branch: develop for all three ecosystems (composer, npm, github-actions) in .github/dependabot.yml, so version updates open against develop correctly.

Security updates are the exception: GitHub always raises Dependabot security updates against the repository's default branch (main), ignoring target-branch — this is not configurable. The default branch is intentionally kept as main (changing it would have repo-wide side effects), so security-update PRs will keep being born against main.

Rule — retarget to develop: whenever a Dependabot PR opens against main (in practice, always a security update), change its base to develop and leave a one-line comment stating the reason (correct flow: every change funnels through develop and ships in the consolidated release PR, never as an out-of-band main commit). Then comment @dependabot rebase so the lockfile diff is recomputed against develop. The PR then rides the normal develop batch with auto-merge like any other.

An agent cannot issue that command, and the failure is silent (#1201). A comment posted through the agent's GitHub tooling has separator characters injected into bot mentions — @dependabot rebase arrived on the PR as ·@·d·ependabot r·ebase, which Dependabot never parses, so nothing happens and nothing reports an error. The PR simply sits. Do not retry the comment; it will be mangled the same way.

What works instead, in this order:

  1. Check whether a rebase is needed at all. It is only needed when the lockfile actually differs between main and develop: git diff origin/main origin/develop -- package-lock.json composer.lock. On the #1201 retarget that diff was empty — develop's batch had never touched the lockfile — so the PR's hunks applied to develop unchanged and there was nothing to recompute.
  2. If the PR is merely behind, update the branch (update_pull_request_branch, the "Update branch" button). That is the whole fix in the common case, and it is what unblocked #1201: it had sat at behind for an hour with every gate already green, and merged 14 minutes after the update.
  3. Only if the lockfile genuinely conflicts does the rebase matter — ask the user to post @dependabot rebase, since @dependabot recreate from a mangled comment is not an option either.

#1201's open question is closed, and by observation rather than by reading the setting: develop requires a branch to be up to date before merging. That is what the hour at behind with every gate green was — a property of the branch, not of Dependabot's flow or of the retarget. #1336 reproduced it on an ordinary two-line docs PR, which merged only after update_pull_request_branch, and its history carries the resulting Merge branch 'develop' into … commit — so the evidence lives in the PR rather than in anybody's memory. The rule that follows applies to every PR here: read mergeable_state, and treat behind as a step to take, never a state to wait out. Step 2 above is therefore the common fix on this repository generally, and behind is the expected state for any PR that develop moved under.

Exception — genuine hotfix: if the security fix is in a runtime dependency (shipped inside the plugin, not dev/CI tooling) and is severe enough to ship to production immediately, treat it as a hotfix (hotfix/* → main) instead of retargeting, then sync develop per "Sync develop with main". Dev/CI-only deps (vitest, @vitest/coverage-v8, undici, js-yaml, PHPStan, etc.) never qualify — they always ride the develop batch.

That exception is unreachable through Dependabot today, and the reason is worth knowing (measured on the #1201 retarget). The plugin ships zero managed third-party runtime dependencies: composer.json's require is php alone, package.json's dependencies is empty, and .distignore drops both /vendor and /node_modules from the zip. So every package Dependabot can see is dev/CI-only, and the decision is mechanical rather than a judgement call — check require / dependencies and the answer falls out.

The flip side is the part that matters: the code that actually ships is invisible to Dependabot. Four third-party bundles are hand-vendored in libs/js/ — altcha-3.2.2.umd.js, thumbmark-1.10.1.umd.js, html2canvas-1.4.1.min.js, jspdf-4.2.1.umd.min.js — all four enqueued at runtime, all four inside the distributable, none in any manifest. A CVE in any of them will never open a PR here; it has to be noticed by a person, so what is running has to be readable off the tree. Tracked in #1202.

That is why the version lives in the filename and the path is built from the constant, never from a literal: FFC_PLUGIN_URL . 'libs/js/jspdf-' . FFC_JSPDF_VERSION . '.umd.min.js'. The two halves then cannot drift — bumping FFC_JSPDF_VERSION without renaming the file 404s the script, which breaks PDF generation loudly. With a literal path the constant is only the ?ver= argument, so the same bump silently serves the old bundle under a new cache key, and nothing anywhere disagrees. html2canvas and jspdf were in exactly that state until #1203 renamed them. One thing that deliberately stays untouched: a vendored bundle's own //# sourceMappingURL= comment still names the upstream filename (thumbmark has had the same dangling comment since it was vendored, and no .map ships). Editing it would break byte-for-byte verification against upstream, which is the reason to vendor at all.

Release PR (develop → main)

When the batch on develop is validated against the testes site and ready to ship to prod:

  1. Open a single PR develop → main.

  2. In that PR (committed onto develop immediately before opening):

    • Bump FFC_VERSION in the three sync sites (ffcertificate.php header, FFC_VERSION constant, readme.txt Stable tag). See "Versioning".
    • Rename the [Unreleased] heading in CHANGELOG.md to [X.Y.Z] (YYYY-MM-DD) and add a fresh empty [Unreleased] heading above it. The release commit reference is appended to that heading later, at step 6 — not here, and not before the tag.
    • Rewrite the == Upgrade Notice == entry in readme.txt to summarise THIS release, replacing the previous one. It is what WordPress prints under the plugin on Dashboard → Updates, so it is the only sentence an operator reads before deciding to apply the update — and the only place a ⚠ breaking change can warn them beforehand. Exactly one entry, naming the version Stable tag declares, at most 300 characters, ending with the section's pointer to CHANGELOG.md for the detail and the issues. One entry because get_latest_release() fetches releases/latest, so an entry for an older version is text no code path can display; 300 because that is the WordPress.org norm, adopted although the directory does not bind us, so publishing there later never means rewriting entries. ReadmeUpgradeNoticeTest enforces all three and calls the updater's own parser rather than mirroring it. This is not a second changelog — that mirror existed, drifted, and was deleted, which the file's own == Changelog == section records; a 300-character summary of the version being offered is a different artefact from a per-release history, and the pointer is what keeps it honest rather than partial.
    • Correct the @since tags the batch wrote against the wrong version. A feature PR cannot know which release it will ship in, so an @since written mid-batch names the version develop was sitting at, not the one the code debuts in. Three classes did this in the 6.25.0 batch (@since 6.24.0 on code that shipped in 6.25.0), while four others had guessed the next number correctly — so the file disagreed with itself. The release PR is the only moment the number is known; git diff origin/main...origin/develop -- '*.php' | grep '^+.*@since' finds them. Nothing else will: no gate reads these, and it was found by accident while looking for something else.
    • Run npm run build:js if any JS/CSS in assets/ changed across the batch and the bundles weren't already rebuilt mid-flight (the "Verify minified assets are up to date" gate would catch this anyway).
  3. Auto-merge SQUASH into main. The squash commit subject should follow main's convention: X.Y.Z — <short summary of the batch>.

  4. Tag the release to publish it. The GitHub Release + distributable zip are automated — .github/workflows/release.yml triggers on pushing a tag matching v*: it validates the tag equals the ffcertificate.php Version: header, builds ffcertificate-X.Y.Z.zip (staged via .distignore), extracts the ## [X.Y.Z] section of CHANGELOG.md as the release notes, and publishes the GitHub Release. So there is no manual "create release" step — after the squash lands on main, tag that commit and push:

    git push origin :refs/tags/vX.Y.Z   # only when re-tagging after a bad tag; skip otherwise
    git tag -d vX.Y.Z                   # ditto
    
    git checkout main && git fetch origin && git pull --ff-only origin main
    test "$(git rev-parse HEAD)" = "$(git rev-parse origin/main)" \
      && grep -q "Version:            X.Y.Z" ffcertificate.php \
      && git tag vX.Y.Z && git push origin vX.Y.Z \
      || echo "STOPPED: HEAD=$(git rev-parse HEAD) main=$(git rev-parse origin/main) / $(grep 'Version:' ffcertificate.php)"

    The two checks are chained with && because a check that only prints does not stop anything, and that is not a hypothetical. The tag has landed on the wrong commit TWICE, and the second time a written-down guard was already in place.

    For 6.23.0 the tag was pushed within a minute of the merge, onto a local main the pull had not yet advanced, so it marked the previous release's squash — Version: 6.22.0 under a tag saying 6.23.0. That produced the guard.

    For 6.24.0 the guard existed and was followed, and the tag still landed on 6.22.0's squash. The pull --ff-only had ABORTED on a dirty package-lock.json, leaving HEAD three releases back; git rev-parse --short HEAD and the grep both printed the wrong values, exactly as designed — and then git tag ran anyway, because the whole recipe was pasted as one block and printing is not refusing. The lesson is not "check": it is that the check must be able to interrupt.

    Two mechanics the recovery taught, both easy to get wrong a second time:

    • Compare against origin/main, never a literal SHA. The first repair attempt hard-coded a 7-character SHA against a git rev-parse --short that returned 8 — --short uses the shortest unambiguous length, which grows with the repository — so a correct HEAD was rejected. The comparison above states the real invariant: this is the commit that is main's tip, and it declares the version being tagged.
    • git fetch --tags does not move a local tag that already exists. After deleting and re-pushing the tag, a local checkout keeps pointing at the bad commit and will happily confirm the wrong answer. Read git ls-remote origin refs/tags/vX.Y.Z, or fetch with --force.

    The workflow's own version check is a backstop and it did fire both times: nothing was published, so there was no wrong zip or Release to retract. But recovering still means deleting a published tag (the first two lines above) rather than never creating a bad one.

    The tagged commit's Version: header must already equal X.Y.Z (true post-bump) or the workflow's sanity check fails. That check is a backstop, not the guard — it fires after the tag is already public. It failing is the good outcome (nothing is published, so there is no wrong zip or Release to retract), but the tag still has to be deleted and re-pushed. The tagged tree will not carry its own CHANGELOG short-SHA suffix, and that is expected — see step 6. Pushing the tag publishes a production release (public zip + Release notes), so an agent surfaces this sequence for the user to run rather than pushing the tag itself — the same production-deploy sign-off that keeps the develop → main PR draft until the user confirms.

    A published Release body is the CHANGELOG section verbatim, and an agent cannot edit one. The Generate release notes from CHANGELOG step awks the ## [X.Y.Z] section out, dropping the heading and stopping at the next ## , so the body and the file say the same thing by construction — until the file is edited afterwards, and then they disagree forever. Nothing reconciles them: release.yml runs once, on the tag. The GitHub MCP server exposes releases read-only (list_releases, get_release_by_tag, get_latest_release) and there is no gh, so editing a published body is the maintainer's to do, like deleting a ref. What an agent can deliver is the exact replacement text — run the workflow's own awk over the edited CHANGELOG.md and the output is byte-for-byte what that release would have published. Measured when #1260 translated the 6.23.0 / 6.24.0 / 6.25.0 sections, whose three Release bodies were left in Portuguese by this limit.

  5. After merge, sync develop with main (see Sync below) so the next batch starts from the bumped baseline. After a release this is a hard reset, not a rebase — the squash already contains every develop commit, so a rebase tries to replay them all and conflicts.

  6. Backfill the release commit reference onto the CHANGELOG heading. Append — `<short-sha>` to the ## [X.Y.Z] heading, matching every other shipped version header. The width is whatever git rev-parse --short returns for the squash commit on main, never a fixed number — this is the same --short property step 4 warns about, seen from the other side: it gives the shortest unambiguous length, which grows with the repository. The shipped headings record that growth, which is how to re-read it: seven characters up to 6.28.4, eight from 6.29.0 on, mixed across 6.28.x because each heading carries what the tool returned the moment its backfill was written. This paragraph said "the 7-char short SHA" for four releases after the answer became eight, nine lines below the sentence that explains why it would. It lands on the post-release develop through a normal PR, and it must come after step 5: the sync is a hard reset to main, so a backfill landing before it is silently discarded — the order #1004 and 037001f followed after 6.20.1 and 6.20.0. The consequence is structural, not a mistake to fix: the tag pushed at step 4 never carries its own suffix, and main acquires it one release later, when this commit rides the next release PR. So a missing suffix on the newest shipped heading is normal; a missing suffix on an older one means the backfill was genuinely skipped (as for 6.11.1 / 6.11.2, fixed retroactively).

Landing the bump on develop (who can do step 2). Step 2's bump commit has to sit on develop before the develop → main PR is opened. A maintainer with direct-push access commits it straight to develop. An agent driving the release does not push to develop directly — not because it cannot (see "Branch naming": the credential has bypass), but because a direct push skips CI on the very commit that sets the version every install will read. So it lands the bump via one dedicated release: X.Y.Z — bump + finalize CHANGELOG PR to develop (auto-merge), then opens the develop → main PR. That dedicated release-bump PR is not what "What not to do" forbids — that prohibition targets feature PRs sneaking a version bump; this is the deliberate release-moment bump, delivered through the only channel an agent has. Either way, the develop → main PR itself stays draft until the user confirms the prod deploy.

Hotfix flow (urgent fix that can't wait for the next release)

When a critical bug needs to ship to prod while develop has un-released commits:

  1. git fetch origin && git checkout -b hotfix/<desc> origin/main.
  2. Apply the fix. Bump FFC_VERSION as a real patch (e.g. 6.7.7 → 6.7.8) — hotfixes consume patch numbers, not the .x.y.z.N cache-bust convention.
  3. PR hotfix/<desc> → main. Auto-merge SQUASH.
  4. Then sync develop with the new main (see below) so develop carries the hotfix and the next release PR doesn't try to "undo" it.

Sync develop with main (post-hotfix or post-release)

git fetch origin
git checkout develop
git rebase origin/main
git push --force-with-lease origin develop

This rewrites develop's SHAs on top of the new main tip. Force-push is permitted on develop by design (the branch protection deliberately omits "Require linear history" and the push restriction) — see "Branch protection" below. If a feature PR was open against develop at the moment of the rebase, the PR author rebases their branch on the new develop tip; this is the cost of keeping develop linear.

Post-release — use reset, not rebase. The rebase above is the post-hotfix recipe, where develop still carries un-released commits to replay on top of the new main. After a release squash-merge, develop was fully consumed by the squash — every develop commit is already inside main's single release commit, so a rebase tries to replay them all and conflicts (typically on CHANGELOG.md). In that case skip the rebase and reset develop straight to main:

git fetch origin
git checkout develop
git reset --hard origin/main
git push --force-with-lease origin develop

Confirm nothing is lost first with git log --oneline origin/main..develop — after a release those are only the pre-squash commits, whose content already lives in main. (This is the trap that bit the 6.15.0 sync.)

Branch protection (develop)

Intentionally lighter rules than main:

  • ✅ Require a pull request before merging (no required reviewers — solo maintainer).
  • ✅ Require status checks to pass before merging — all gating jobs listed under "CI gates".
  • ✅ Require branches to be up to date before merging (strict checks). Observed, not read — see the limit below.
  • ❌ Require linear history — left off so the rebase workflow above doesn't need admin bypass.
  • ❌ Restrict who can push to matching branches — leaving force-push permitted is what makes the rebase sync above mechanical.
  • ❌ Require deployments to succeed — deploy-develop.yml runs after merge, not as a merge gate.

Reasoning: develop is single-maintainer integration territory, not a shared production branch. Stronger protection here would force admin bypass for routine syncs and provide negligible safety benefit. The strict-checks line is the one rule here that is not lighter, and it is the one this list did not carry until #1337 — whether it was chosen or inherited is not recorded anywhere a reader can check. Its cost is concrete either way: every branch update re-runs every gate.

An agent cannot read this configuration, so the list above is a record and not a reading. The GitHub MCP server exposes no branch-protection or ruleset endpoint — list_branches answers protected: true for develop and for main, and nothing further — and there is no gh. Each line here is therefore either the maintainer's own record of the UI or, for strict checks, an inference from behaviour that a PR's own history preserves. Treat it the way "Branch naming" treats the push rules: check before relying on it, because a written-down wall that nobody can re-read is exactly how that section came to claim a wall that was not there.

Which UI page, though, is itself uncertain, and the evidence points away from the obvious one. This section used to open "Configured in Settings → Branches". But the post-6.25.0 sync force-push reported Bypassed rule violations for refs/heads/develop — rule violations is ruleset wording, not classic branch-protection wording — so at least part of what governs develop lives in Settings → Rules → Rulesets. The two mechanisms stack and both can require status checks, so a setting that seems missing from one page may simply be on the other. Check both before concluding a rule is absent.

Deploy to testes

.github/workflows/deploy-develop.yml runs on every push to develop and rsyncs the working tree to the testes server. Required GitHub secrets (Settings → Secrets and variables → Actions):

Secret Example Notes
TESTES_SSH_HOST ssh.testes.example.com or 185.239.210.8 DNS or IP of the testes host. Hostname only, no port, no protocol prefix.
TESTES_SSH_USER wp-deploy Account with write access to the plugin dir
TESTES_SSH_KEY -----BEGIN OPENSSH PRIVATE KEY-----… Private half of a dedicated keypair; public half goes in ~/.ssh/authorized_keys on the testes host. Must have no passphrase — generate with ssh-keygen -t ed25519 -N "" -f <path>. GitHub Actions cannot enter passphrases interactively; a passphrase-protected key surfaces as Permission denied (publickey,password) in the rsync step, indistinguishable from a wrong key.
TESTES_SSH_PORT 65002 Optional. Defaults to 22. Managed hosting (Hostinger, KingHost, Locaweb) usually exposes SSH on a high port — set this when so.
TESTES_REMOTE_PATH /var/www/testes/wp-content/plugins/ffcertificate Absolute path; no trailing slash

The rsync uses --delete, so anything in the remote path that isn't in the develop working tree is removed on each deploy. The workflow excludes .git/, .github/, vendor/, node_modules/, tests/, and dev tooling (PHPStan, PHPUnit, PHPCS configs) — those don't belong in a runtime plugin dir.

The testes server should have SCRIPT_DEBUG=true in wp-config.php so non-minified assets load and ?ver=… cache aggressiveness stays low while iterating.

Post-deploy smoke (alarm, not a gate)

After the rsync, deploy-develop.yml scp's .github/scripts/testes-smoke.php to the host, runs it over SSH and deletes it. It boots the real WordPress and asserts three things, then reports the scheduled ffc* crons without judging them (which crons should exist depends on the per-module toggles, so a strict expectation would fail on a deliberate configuration):

  1. The plugin is active — a deploy that lands files onto a deactivated plugin looks healthy from the filesystem and does nothing.
  2. The host's live FFC_VERSION equals the version in the commit being deployed. This is the #628 check: the rsync backoff was once shorter than a managed-hosting restart, so every attempt fell inside the same outage and the site quietly stayed on a stale version while the workflow reported success.
  3. Every table uninstall.php knows about exists. The expected list is parsed out of uninstall.php rather than duplicated — that file is already obliged to know every table, so a new one is covered the moment it is added there. It is read as text and never included (including it would run the uninstaller). If the parse yields nothing the smoke fails loudly instead of passing on an empty list.

This exists because the host is the only place with the provider's real MariaDB, PHP SAPI and wp-config — the stack that produced the dbDelta failures CI cannot reproduce (#358 backtick comments read as columns, #822 a migration pointing at a table that never existed, 6.0.1 COMMENT clauses leaving four recruitment tables uncreated).

Two constraints to keep in mind before extending it. It can never be a merge gate — this workflow fires on push to develop, so the merge already happened; gating belongs in CI, against a MySQL service. And it runs against an established install, so the table check proves the tables are there, not that a fresh activation would create them; catching a regression in the CREATE of an existing table needs a throwaway WordPress install (not merely an empty database — Activator::activate() reads and writes options). That is the fresh-install CI job (see "Quality gates and testing" → "Fresh-install gate"), delivered in #994.

The step blocks: a red smoke is a red deploy job. It ran under continue-on-error: true while it proved itself against the host's PHP CLI, and #992 set the bar at 10 consecutive green smokes — an evidence gate, not the release-cycle deprecation, which exists for surfaces whose consumers a code scan cannot see; here every deploy is evidence. Runs 525-537 delivered 13, verified from the log output rather than the step conclusion, and the flag came off in #1007. Two exits stay green on purpose. A timeout (124) is a warning, not a failure: the rsync has already succeeded, so a host that does not answer in 120s is latency, not a plugin defect, and reddening deploys for it is how an alarm becomes noise people learn to skip. And a host whose CLI is older than the plugin's PHP floor skips with an explanation rather than failing forever — point the workflow at the host's newer binary (often php83) to enable it. Everything else — a stale version, a missing table, a fatal in the plugin — now fails the job. The reason the evidence had to come from the log is worth keeping: while the flag was on, a failing smoke still showed the deploy job as green, so the job conclusion proved nothing and only the step's own SMOKE PASSED line did.

Versioning

Three places carry the plugin version and must stay in sync:

  1. ffcertificate.php plugin header — * Version: X.Y.Z. Parsed by WordPress core BEFORE PHP runs, so it must be a literal string.
  2. ffcertificate.php PHP constant — define( 'FFC_VERSION', 'X.Y.Z' ). Source of ?ver=… on every wp_enqueue_* call.
  3. readme.txt Stable tag: X.Y.Z. Parsed by WordPress.org before PHP runs; also a literal string.

When changing the version, update all three in the same commit.

Patch vs. cache-bust-only releases

  • A "real" patch release (any source-code change) consumes the next patch number: 6.6.2 → 6.6.3.
  • A cache-bust-only release (no functional change — exists purely to rotate the ?ver=… asset cache key after a prior PR shipped an updated .min.js / .min.css without bumping the version) uses a 4th segment appended to the prior version: 6.6.2 → 6.6.2.1. The next cache-bust sibling of the same minor would be 6.6.2.2, and so on. WordPress's version_compare() and the plugin update flow both handle 4-segment versions without special-casing.
  • Reason for the convention: a cache-bust release carries no new user-visible behavior, only a key rotation. Burning a real patch number on it would imply meaningful changes that aren't there.

When to bump

The trigger has not changed — bundled-asset changes still rotate the cache key. What changed with the develop branch workflow is where the bump lands:

  • PRs targeting main (release PR develop → main, hotfix PR hotfix/* → main): bump FFC_VERSION in the same PR. The release PR consolidates every assets/**/*.min.js, assets/**/*.min.css, templates/**.php, and languages/*.l10n.php / .mo change from the develop batch under one version. Hotfix PRs bump their own patch number.
  • PRs targeting develop: do not bump. Develop sits at the last released version (the cache key on the testes domain stays stable across the batch), and the testes site sidesteps cache aggressiveness via SCRIPT_DEBUG=true. Bumping per-PR on develop would consume version numbers that have no production analog.

The "Verify minified assets are up to date" CI job catches build freshness on both bases but does NOT enforce the version bump — that's still a human discipline on the release PR.

CHANGELOG conventions

CHANGELOG.md follows Keep a Changelog. Per change:

  • One entry under the top [Unreleased] section, grouped by heading (Added / Changed / Fixed / Security / Removed / Deprecated). Entries stay in [Unreleased] across PRs until the release PR renames the heading (see "Release PR").

  • Always cite the issue/PR ((#NNN) / #NNN). Every [Unreleased] bullet must carry a reference — the linked PR holds the granular detail.

  • Append the bullet at the END of its heading's list, never at a fixed point inside it — this removes the DECISION, not the conflict. Two branches open against develop at the same time both edit [Unreleased], and appending does not stop them colliding: two appends land on the same line and git reports it exactly as an insert would. What appending buys is that the resolution is always keep both, in either order, with nothing to judge — where inserting before a chosen existing bullet makes the conflict about where the new line goes relative to someone else's. Measured across the #1260 arc, which ran six slices: five conflicts, every one resolved by keeping both. The first version of this bullet claimed appending prevents the conflict; it was written between the two slices that then collided by appending, which is how the claim survived being false. Order of arrival has no value here anyway, because the release PR condenses the section by story.

  • No internal roadmap codenames in the prose — no "Sprint N", "phase N", or letter-codes (A6, B3, E5, …). Describe the change itself and keep entries concise (one tight paragraph, not a wall of class-by-class text).

  • Ordinary words that happen to look like codes — "A4" (paper size), "four-phase flow" (literal steps) — are fine; the rule targets roadmap taxonomy only.

  • Aim for ≤300 characters per bullet. A soft target, not a CI gate — a bullet may exceed it when the detail genuinely earns the length (a breaking-change banner, a subtle regression), but prefer trimming first: the linked PR holds the granular detail, so the CHANGELOG line only needs the what + the reference. When a batch accumulates several long or near-duplicate bullets, condense before the release PR (the precedent that set this: the #772-era [Unreleased] condensation).

    6.25.0 measured what "condense" means at scale, because the rule said to do it without saying how much or how to notice it was time. That [Unreleased] reached 85 bullets and 35 KB, median 404 characters, and went out at 37 bullets and 10.6 KB, median 287 — 70% less text, and still longer than 6.24.0's, which shipped 40 bullets for a smaller batch. Two signals said it was overdue and both are cheap to check: the median bullet length sitting well above the 300 target, and duplicate headings (### Added ×2, ### Changed ×2) — the tell that batches were appended without anyone re-reading the section as a whole. Condense by story, not by PR: nine bullets citing one issue became four, and two CSS issues dissolved into the themes they belonged to. Then prove nothing was lost by measurement rather than memory — every issue that delivered work must still be cited (41 of them, checked programmatically); references that leave are only the precedents quoted in prose. This matters more than it looks: release.yml extracts that section verbatim as the public Release notes, so the section is not an internal file, it is what the user reads.

What not to do

  • Do not amend or rewrite published commits on main. On develop, force-push is permitted only for the documented sync-with-main rebase ("Develop branch workflow" → Sync) — never to rewrite arbitrary history.
  • Do not skip hooks (--no-verify) or signing.
  • Do not bypass the coverage floor — bump it forward or restore the lost coverage. Never lower it — the sole exception is an honest re-measure after deleting covered product code (never tests); see "CI gates".
  • Do not add new untested code paths in a coverage-aware PR; either cover them in the same PR or document the deferral.
  • Do not target main directly from a feature PR. The only PRs that base on main are the release PR (develop → main) and hotfix PRs (hotfix/* → main).
  • Do not bump FFC_VERSION in a feature PR that targets develop — the bump belongs to the release. (The lone exception is the dedicated release: X.Y.Z bump PR an agent uses to land the bump on develop when it can't direct-push; see "Release PR" → "Landing the bump on develop".)

2. Quality gates and testing

CI gates (all gating, enforced on main and develop)

  • PHP: PHPStan (level 8) · WPCS · PHPUnit (8.3/8.4) · Coverage ≥ floor (clover, env COVERAGE_FLOOR_LINES in .github/workflows/ci.yml).
  • JS / CSS: ESLint (zero-error) · Stylelint (zero-error) · Vitest + coverage ≥ floor (env JS_COVERAGE_FLOOR_LINES in .github/workflows/lint.yml).
  • Misc: CodeQL (javascript) · Composer audit · Review dependency changes · Verify minified assets are up to date · Fresh install (PHP 8.3) — see below.

The same gates run on PRs to develop and on the release PR develop → main. Develop must stay deployable to the testes site, so we don't relax gates there — a green develop is the precondition for opening the next release PR.

The coverage floors are ratcheted upward in the PR that delivers the gain — never lowered. The comment block above each *_FLOOR_LINES keeps the audit trail.

One exception — code deletion. Removing well-covered product code (never tests) can legitimately drop the line-% because high-coverage lines left the denominator — that is not a regression to restore. When a deletion PR lowers the measured coverage, an honest re-measure of the floor down to the new figure is allowed, provided the *_FLOOR_LINES comment block records the deleting PR and the new baseline. This is the only case the floor may move down; deleting or weakening tests to relax it never qualifies.

Acceptable floor buffer: when bumping, the floor may sit up to 5 percentage points below the freshly measured coverage. A buffer of ≤5pp is acceptable for both JS (JS_COVERAGE_FLOOR_LINES) and PHP (COVERAGE_FLOOR_LINES) — it absorbs v8/clover run-to-run jitter so the gate doesn't flake on fractional swings, without forcing the floor to chase every decimal. So: still ratchet up when a PR delivers a real gain, but leave no more than ~5pp on the table, and never set the floor above the lowest run you've actually observed.

The guards

Every guard below runs inside the PHPUnit gate above, except the two that need a database and live in the fresh-install job. Each is here because a defect had already shipped — none was written on spec, and the paragraph under each name records what it cost to find, because that is the part a rewrite loses.

Two properties they share, and both are load-bearing. Every one carries a self-check that fails when its own scan collapses — an empty result must never read as clean (the #1071 / #1094 rule). And every one sees presence, never truth: a wrong reason in an allowlist passes, an assertion against the mock that supplies the value passes. Crossing the boundary a guard mocks is what the fresh-install job does for schema, and what a render does for CSS.

Whether a seam is worth a guard at all is decided by what its contract is made of, not by how much the seam hurts (#1525). Three defects found by use rather than by a gate — a dead admin_post handler, a submit marker the browser never posted, a queue that hid a number from the panel that corrects it — looked like one class needing one sweep. They are not: only one crossed PHP and JavaScript, one crossed our code and WordPress core, and one never left PHP. What they share is that each side was individually correct and each side's test mocked the other, which is a property of tests, not of a language boundary.

So the line is the contract, and it has two sides:

  • A number or a name declared on both ends is cheap to guard, and worth it. An accepted_args is an integer checkable against a signature; an AJAX action and an option key are literals at both ends. ZeroArgHookArityTest and AjaxWiringTest are that shape, and both cost little.
  • The SHAPE OF AN OBJECT crossing two languages does not get trustworthy cheaply, and a scan that is not trustworthy is worse than none. The localized-payload seam was measured and came back with no live defect, after four calibration rounds each wrong differently — an object that is also the JS namespace, nesting on both sides, a prefix matching inside a longer name, and finally JS scope, which regex cannot reach because the same local is rebound to a jQuery object in another function. A gate at that false-positive rate gets allowlisted into silence, which is the failure this file already names when it explains why a post-deploy timeout stays green.

The data-ffc-* seam measured the same way: every apparent dead listener was emitted through a data array whose renderer prefixes data-, so the literal exists nowhere — the trap CssClassEmitters already records. Do not re-run either census on a hunch; #1525 holds the method and the result.

A self-check is worth what it is tight to, and a floor loosens on its own (#1428). The rule above is about a scan reading nothing; a scan reading half is the harder case, and measured across the 30 guards that walk a directory, seven could not tell a partial collapse from a clean tree — assertGreaterThan( 0, … ) over annotations numbering in the hundreds, 50 over 128 views, 4 over ten enqueues, assertNotEmpty over three files. Nobody loosened any of those: a loose floor is what a tight floor becomes. > 100 was tight against 120 files and is slack against 250, and the population grows underneath without anyone touching the guard — CssNamespaceAnchorTest's 25 of 28 sheets detects today and will not the day a fortieth lands. Four shapes hold instead. The order below is the order to reach for them, and the first two are not competing — which one applies is decided by what the guard already has:

  1. PREFERRED, when a non-empty frozen register already exists — an exact comparison against it. assertSame against a baseline tolerates nothing, costs no new code, and is what most of the guards already do: it is why 19 of the 30 killed the mutation. AssertionCoverageTest is the strongest version, because a shrink-only ratchet is itself the detector — a scan that loses files loses the baselined entry from its findings, and the "a baselined test now verifies something" direction fires. The condition is the register being non-empty: a baseline driven to zero gives this up and owes an explicit self-check in exchange, which is what CssNamespaceAnchorTest has.
  2. PREFERRED, when there is no register — an independent recount, taken by a different traversal from the one under test: scandir against RecursiveIteratorIterator, or scandir against glob. ActivatorSqlTest has held this shape all along and its own comment states the rule — compare counts rather than asserting a floor. It is the only shape that never needs maintenance, because adding a file moves both sides. Its cost is real: where the scan filters, the recount repeats the filter, and the two can drift. Prefer it anyway — a drifting recount FAILS, while a decaying floor goes quiet, and the failure mode that speaks is the one to buy.
  3. ACCEPTABLE, only when repeating the filter would be dishonest — a floor stated as a share of an independently counted population. PhpcsSuppressionTest skips SKIP_DIRS and includes/libraries/, so an exact recount would mean copying both rules into the self-check. A share survives population growth, which an absolute floor does not — but it trades growing slack for constant slack: 90% of 28 sheets still tolerates losing two, and losing two quietly is exactly what #1423 was. Reach for it to avoid duplicating a filter, never to soften a comparison that could have been exact.
  4. LAST RESORT — naming a single deep member. PhpcsSuppressionTest names uninstall.php, the last of its 591 sorted entries. Honest as long as the file named is one the scan can least afford to miss; when it legitimately leaves, the guard fails and asks to be re-pointed, which is maintenance the guard owes rather than a defect. It proves the read reached that one place and says nothing about the rest.

NEVER — a bare floor. An absolute one decays as the population grows past it, which is how all seven defects were born. A share derived from the scan's own count is worse than useless: it detects nothing at all, because halving the scan halves numerator and denominator together and the ratio holds at one.

An audit that halves a population to test any of this must prove the population shrank. Three verdicts in the #1428 measurement were mutations that mutated nothing — a one-element glob and a one-entry ratchet, where half is the identity — and a two-state pass/fail report counted both as verified. The same census also read a tearDown cleanup as a guard, because it defined one as "walks a directory and reads files". Re-run it as a one-off when the guard population grows; never as a gate, for the reason the UnusedFunctionParameter verdict gives below.

Module-boundary guard (#563 B3)

tests/Unit/ModuleBoundaryTest.php freezes the cross-module dependency graph of includes/ (a "module" = the first namespace segment after FreeFormCertificate\; an edge = module A referencing FreeFormCertificate\B\…) against the committed baseline tests/fixtures/module-boundary-baseline.php. It runs in the normal PHPUnit gate. The graph is a ratchet that can only shrink: a new edge fails (new cross-module coupling — justify it or route through a facade); a removed edge also fails (coupling eliminated — lock the win in). After an intentional change, regenerate + review the diff: FFC_UPDATE_BOUNDARY_BASELINE=1 vendor/bin/phpunit --filter ModuleBoundary. Never regenerate just to make a red guard green without understanding the new edge.

Fresh-install gate (#994)

.github/workflows/ci.yml → job fresh-install installs a throwaway WordPress onto a MariaDB service, activates the plugin into an empty database, and runs .github/scripts/fresh-install-check.php. It is the only gate that sees a fresh activation: the post-deploy smoke runs against the established testes install, where every table already exists from an earlier release, so a CREATE that stopped working is invisible there. That is the 6.0.1 class — and it was live again when this job was written (ffc_self_scheduling_calendars was never created on any fresh install, because dbDelta() splits its input on ; and an SQL comment inside the statement carried one).

A fresh install also makes a check possible that no established install can make honestly: the ffc_* tables and options in the database are exactly what this activation created, with no legacy residue to explain away. So the comparison against uninstall.php runs both ways — every declared object must exist, and every existing object must be declared. The second direction found two tables and 21 activation-written options that deletion had never removed; most are written through a constant or a variable, so no static scan could have seen them. The job then deletes the plugin with the Danger Zone opt-in on and asserts the footprint is zero, which makes uninstall.php an enforced manifest rather than a list nothing checks.

Three things to know before touching it. uninstall.php is the single manifest — .github/scripts/ffc-uninstall-manifest.php parses it as text (never includes it) and both this job and the post-deploy smoke read it, so the two can't disagree about the footprint; adding a table or option to uninstall.php is what covers it everywhere. The MariaDB service image is a deliberate pin — the point is to reproduce the provider's dbDelta behaviour, so the smoke prints the host's real VERSION() and @@sql_mode on every deploy and the pin moves when those drift. It was pinned to 10.11 on a guess and corrected to 11.8 the moment the first smoke reported the host's actual 11.8.8-MariaDB-log (whose sql_mode also lacks ERROR_FOR_DIVISION_BY_ZERO) — read the smoke, don't pick a familiar LTS. It is not a substitute for the smoke: the runner's MariaDB is not the host's, and only the host has the provider's PHP SAPI and wp-config. The cheap half of the same defect class runs with no database at all in tests/Unit/ActivatorSqlTest.php, which enforces two dbDelta rules on every CREATE TABLE literal: no semicolon before the final one (it splits its input on ;, truncating the statement), and every line a column or key definition — a blank line or an SQL comment is read as a column and becomes a malformed ALTER on every run against an existing table. Measured against a real MariaDB, the three self-scheduling tables produced 41 database errors before that cleanup and 1 after (#997); the survivor is a Duplicate key name 'validation_code', where the CREATE declares a plain KEY that ensure_unique_validation_code_index() later converts to UNIQUE. All of it was latent — every affected activator guards its dbDelta call on table_exists() — and would fire the day someone adds a column the normal WordPress way.

dbDelta idempotence gate (#1087)

A step in the fresh-install job replays every CREATE TABLE through dbDelta() against the table it just created from that very statement. dbDelta ALTERs away the difference between a statement and what the server stored, so when the two disagree in a way the statement can never satisfy it re-ALTERs on every run that reaches it — the #997 class, which had been measured exactly once, by hand. Nothing else sees it: ActivatorSqlTest reads the statement as text with no database, the rest of fresh-install only ever watches the CREATE succeed, and the post-deploy smoke asserts the tables exist, not that the schema still matches. The statements come from .github/scripts/ffc-create-statements.php, shared with ActivatorSqlTest so the two cannot measure different sets, and the gate aborts on an empty scan, a count that diverges from a wider net, or a table name it cannot resolve — it never counts as clean what it did not look at.

It blocks at zero. The first measurement found 14 statements drifting and all 14 were fixed, so the frozen baseline that carried them in the meantime is gone — a change the gate reports is a change to fix, never a line to add somewhere. Diagnosis needs the database, which only that job has, so on failure the gate dumps SHOW CREATE TABLE for the statements whose disagreement is not a type spelling; there is no way to reproduce it locally.

Three things the first measurement taught, all of them the opposite of what they look like. Write the display width. int unsigned looks cleaner and is wrong here: MariaDB stores int(10) unsigned and compares textually, while WP core's dbDelta ignores a width-only difference on MySQL 8.0.17+ and explicitly not on MariaDB. So int(10) is correct on both servers, and it is what WP core writes in its own schema — do not "tidy" it away. json is not stored by MariaDB, which implements it as LONGTEXT plus a CHECK, so a json column can never match there. The ten such columns are declared longtext (#1087 passo 9) because nothing in the plugin uses a native JSON function — every read is json_decode in PHP, and PreflightStatsService already aggregates in memory rather than depend on JSON_EXTRACT; declaring the intent bought validation the code never relied on, at the cost of a statement that could never be honest. Declare the index the code actually creates: Activator::upgrade_auth_code_unique_constraints() converts auth_code to UNIQUE INDEX uq_auth_code on two tables, so a declared plain KEY auth_code never survives — the same shape as validation_code in the self-scheduling tables.

Schema-agreement guard (#1087)

tests/Unit/SchemaAgreementTest.php covers the two shapes schema duplication takes here, without a database. Within one file: a CREATE TABLE and an add_columns_if_missing() call must name the same columns — a fresh install gets the first, an upgraded install the second, and ffc_submissions was born with 7 of its 25 columns until #1091 because nothing compared them. Across files: where a table is declared by more than one class — the activators were written after the migrations that first created those tables, and neither side was retired — every declaration must name the same columns and the same key names. The count is deliberately not stated here. It shrinks as duplication is retired, and it had already gone stale in this paragraph: three tables at ×3, ×3 and ×2 when it was written, two tables at ×2 each when it was re-measured. The enumeration lives in the guard's own docblock, beside the scan that produces it, with the reading written down so it can be re-run rather than believed. That second direction found the third occurrence of the #1091 class — UserDashboardActivator declared 12 of ffc_custom_fields's 17 columns, and a fresh install came out right only because a one-shot migration ran later in the same activation. It also found why installs carry two indexes on auth_code and two on magic_token: both paths indexed the same columns under different names. It compares names, not types (the sources spell types differently by construction) and reads the statements through .github/scripts/ffc-create-statements.php, the same parser the dbDelta gate uses.

Vacuous-test ratchet (#997)

tests/Unit/AssertionCoverageTest.php fails when a new test method verifies nothing — no Mockery/Brain\Monkey expectation and no assertion beyond a literal assertTrue( true ). The register is frozen in tests/fixtures/vacuous-tests-baseline.php, a ratchet that can only shrink: a new entry fails (write a real assertion), and a removed one also fails (a test was fixed — lock the win in). It is closed: 1 entry, not the 108 it opened with — #997 stage 2 fixed the 13 in schema/activation and security, and #1030 (6.22.0) took the remaining 95 down to one. Most of those 94 drove a nonce or capability check and then ended in assertTrue( true ), so deleting the guard they existed for left them green — which is also how each fix was verified: delete the check, confirm the test now fails. That is the method to reuse, because a test that asserts the right thing for the wrong reason passes any static checker. The survivor is kept deliberately, with the reason written in the test: PluginActivationSmokeTest::test_foreign_key_migration_composes_when_version_stale has no stronger observable without a real database — two were tried, and neither distinguishes the branches — so what it asserts is composition without a fatal, and proving the ALTERs belongs to the fresh-install job that has the database. Regenerate after an intentional change with FFC_UPDATE_VACUOUS_BASELINE=1 vendor/bin/phpunit --filter AssertionCoverage and review the diff.

Two things to know. A checker that only looks for $this->assert* reports most of this suite as vacuous — Functions\expect( 'wp_mail' )->once() and shouldNotReceive() are real assertions, so the guard's expectation list must stay in sync with the idioms in use. It sees presence, never strength: a test asserting the wrong thing, or asserting against a mock that supplies the very value under test, passes it. That class is only reachable by reading the test or by crossing the boundary it mocks — which is what the fresh-install gate does for schema. The baseline is a debt register that reached zero-plus-one, not a target to grow: an entry added to it is a decision to be argued in the test, the way the survivor argues its own.

AJAX wiring guard

tests/Unit/AjaxWiringTest.php cross-checks, statically and without WordPress, that every registered wp_ajax_* handler is referenced somewhere other than its own registrar, that every action the client asks for is registered (on wp_ajax_*, admin_post_* or admin_action_*), and that every get_option( 'ffc_*' ) key has a write path. It targets the "built but never wired" defect class — the ten caller-less handlers swept in #935, and the #936 read of ffc_cleanup_days, an option nothing ever wrote, which left submission auto-delete dormant for every install.

Two things to know before touching it. Registration has two idioms — the literal add_action( 'wp_ajax_ffc_x', … ) and the endpoint-class add_action( 'wp_ajax_' . self::AJAX_ACTION, … ); both are resolved, and a checker that knows only the first reports every *AjaxEndpoint action as unregistered. The client names an action in at least six shapes — a FFC.request argument, an action: key, an 'action=…' query fragment, a FormData.append pair, a hidden input's value, and a wp_localize_script payload carrying the constant (never a literal). Direction A therefore matches the bare token, deliberately loosely: matching shapes exactly reported all six as orphans. Direction B matches precise request positions, because it compares against the registered set and a loose match there would flag nonce names and capability slugs.

It proves only that both ends of the wire exist and agree on a name — not that the handler is reached from the right screen, that its button renders, or that the request succeeds. Those need a real WordPress and a browser. Exceptions go in the KNOWN_* allowlists with the reason inline; an action that merely lost its caller does not belong there — delete the handler.

CSS namespace ratchet (#1152)

tests/Unit/CssNamespaceAnchorTest.php freezes the selectors in assets/css/*.css that name nothing the plugin owns — the baseline is now empty (it opened at 60; #1170 took it to 51 by prefixing the ones that were ours, #1184 to 3 by giving the rest a page anchor, and #1202 item 1 to zero by renaming the last three), a ratchet that only shrinks: a new anchorless selector fails, and a baselined one that gained an anchor also fails (drop it from the list to lock the win in). :root is the one allowlisted entry, with its reason: it is how a custom property is declared, and every property under it is --ffc-*.

The last three were the case where anchoring is the wrong fix, and that generalises. They were the user dashboard's #tab-* ids, and an id is unique in the document: anchoring inside a container leaves the generic name exposed, so a theme declaring #tab-profile does not merely repaint — it breaks getElementById, the tabs' aria-controls and the event delegation. The fix is the rename, and the name was already in the repository: the form editor's tabbed container emits ffc-tabnav-<key> / ffc-tabpanel-<key>, and the dashboard already had the nav half right (ffc-tab-<slug>) with only the panel unprefixed. So look for the convention before inventing one — it was found by a proof harness colliding with it, not by a search. Zero is not the end of the guard: what charges it is test_no_stylesheet_publishes_a_new_anchorless_selector(), which scans all 28 sheets every run; an entry added back to the baseline is a decision to defend, not a line to add.

It is the only guard in the theme arc that is prevention rather than a defect already shipped — nothing is broken, no collision observed. A rule like .button::before { content: '\f123' } reaches every button on the screen, not only ours; the radius is small today because the sheet loads on two pages, and that is enqueue luck, not design. AudienceAdminPage::print_menu_separator_css() exists precisely because one rule already had to leave ffc-audience-admin.css when its target started appearing on screens where that sheet does not load.

Four things the measurement taught. Decide the anchor rule before measuring, and mind the token boundary: a[href^="#ffc-separator-"] is anchored — it cannot match anything that is not ours — but the character before ffc there is #, so (?:^|[-_])ffc[-_] drops seven selectors and reports 67 instead of 60; the number then describes the pattern rather than the CSS. Count selectors, not classes: in .appointment-status.status-pending the second class never appears alone, so the compound exposes one name — counting classes reported eleven where there was one. The scanner has to be quote-aware and skip @keyframes: img[src^="data:image/png;base64"] carries a ; that cuts the selector in half otherwise, and 0% / from are steps, not selectors. And the issue's own recorded figure was 40, not 60 — it had missed .button*, .card/code, .form-table* and the #tab-* ids, so re-measure rather than trust the number written down.

What it does not see: inline CSS printed from PHP (print_menu_separator_css(), the appointment receipt) and style="" attributes — it reads the sheets only. Nor duplication: the same scan found 19 ffc-* classes declared bare in more than one sheet (.ffc-status-badge in five), which is a component without a single owner, not a namespace problem, and is tracked apart in #1162.

Page-scope guard (#1184)

tests/Unit/AdminPageScopeTest.php fails when a div.wrap the plugin prints does not carry ffc-admin-page plus exactly one registered ffc-page-<slug>, when a registered screen class is emitted nowhere, or when a NESTED exception stops matching a real wrap. It is the markup half of the #1152 fix — the anchor those 41 baselined selectors need in order to be rewritten — and it deliberately proves only that the anchor is there, never that a rule reads it; that is CssNamespaceAnchorTest's job. The convention, the screen map and the three things the measurement found are in "Stylesheets and theme" → "Page scope on the admin wrap".

Component-ownership guard (#1162)

tests/Unit/StylesheetOwnershipTest.php fails when a class is declared bare — .ffc-x { }, no ancestor and no second class — in two sheets with no dependency edge between them, in either direction, direct or transitive. Without that edge the winner is enqueue order, which is the order modules are wired in Loader: moving two bootstrap lines repaints a screen. Pairs that cannot coexist on one screen (admin × frontend, distinct admin screens) go in ALLOWED with the reason, because that impossibility lives in the enqueue gates, which this scan does not read.

It was not prevention — the defect was live. .ffc-status-cancelled was declared by ffc-calendar-admin.css (red, --ffc-danger-*) and by ffc-audience-admin.css (amber, --ffc-warning-*). Appointments is a submenu of ffc-scheduling, so the audience sheet's gate (strpos( $hook, 'ffc-scheduling' )) matches there too and both load; AudienceLoader is wired after SelfSchedulingLoader, so amber won. In the Status column "Cancelled" rendered amber while its four siblings rendered as the module's own sheet asked. Both families were given their own names — ffc-appointment-status-* and ffc-audience-status-* — in the #1151 / #1154 mould.

Three things the measurement taught. ffc-admin-submissions.css loads on every ?page=ffc-* screen, not only the submissions one: its gate is is_ffc_page(), which matches any ffc- menu, while the enqueuer's own docblock says "submissions page". That is why its .ffc-status-badge collided with recruitment's and reregistration's, which were the two entries the baseline opened with; both were resolved by the renames in #1183, so BASELINE is now empty and the gate blocks at zero. The result is a MERGE, not "the later one wins": the audience badge was rendering text-transform and letter-spacing that only the submissions sheet declares, and white-space: nowrap that only ffc-common.css declares — three properties absent from the component's own sheet, so neither "last wins" nor "first applies" describes what shipped. The rename declares the whole shape on purpose, which keeps the render identical and makes it intentional. And one handle can be enqueued with different dependency lists at different sites (ffc-admin-settings is one; WordPress keeps whichever registered first) — the scan unions them deliberately, because what matters here is whether the edge is declared anywhere.

Two mechanics for anyone extending it. The extractor must read top-level arguments, not explode( ',', … ): plugins_url( "…", dirname( __DIR__, 1 ) ) carries an inner comma that cuts the list in the wrong place and drops ffc-calendar-admin.css from the graph entirely — and it resolves self::HANDLE_CSS against the file's own constants, without which ffc-recruitment-admin.css disappears the same way. test_every_stylesheet_is_reachable_through_an_enqueue() is what stops either from passing as "clean"; ffc-appointment-cancellation.css is the one sheet with no handle, because its handler prints the <link>s itself in explicit order. The CSS parser is shared with the #1152 guard through tests/Support/CssSelectors.php, for the same reason .github/scripts/ffc-create-statements.php is shared: two guards measuring the same sheets must not disagree about what a selector is.

Suppression guards (#1027, #1035)

Two guards cover the annotations that switch a gate off. tests/Unit/PhpstanSuppressionTest.php fails on a @phpstan-ignore-next-line (which suppresses every error on the line — the identifier after it is only a comment; only @phpstan-ignore <id> filters), on one naming no identifier, and on one with no reason. tests/Unit/PhpcsSuppressionTest.php fails on a bare annotation, one with no -- reason (or one under 15 characters — a floor against "ok"/"see above.", never a quality bar), and on the two structural defects below.

PHPCS annotations are a flat on/off switch per sniff, not a stack. A phpcs:enable X anywhere in a file turns X back on for the rest of it no matter who turned it off — so an inner disable/enable pair silently ends an enclosing file-level disable at the enable, leaving the tail of the file uncovered while the file-level comment still claims otherwise. This was live in four classes and nothing was red, because each tail stayed covered by whatever the inner pairs covered; it is only reachable by reading, or by the guard that now checks it. When an inner pair names a sniff an outer disable already covers, drop it from the inner pair.

WordPress.DB.DirectDatabaseQuery is disabled at file level in ~50 classes, and that is scope-gated. The repositories, activators and migrations query the plugin's own ffc_* tables, for which WordPress exposes no API, so the justification is a property of the class and one file-level disable beats one annotation per statement (#1035 collapsed 375 of them into 49). The disable is only honest while every table the file touches is ffc_* — test_file_level_direct_query_disables_only_cover_plugin_tables() enforces exactly that, so adding a wp_posts/wp_users/wp_usermeta query to one of those files fails CI. Fix it by using the WP API or by narrowing to per-line annotations; the CORE_TABLE_EXCEPTIONS allowlist is for a reference that is unavoidable (MigrationForeignKeys must name wp_users — its foreign keys point there) and takes the reason inline.

Both guards see presence, never truth, and neither sees deadness. A wrong reason passes; a suppression a refactor made inert passes. Deadness is measured, not asserted, and the measurement is deliberately out of CI (it costs a full re-run over a rewritten tree): neutralise every annotation in place — rewrite the sniff it names to one that cannot fire, never rename the phpcs:/@phpstan- token itself, which turns the directive into an ordinary comment and trips Squiz.Commenting.InlineComment on hundreds of lines that were fine — then re-run the tool and map the violations that surface back to the annotations covering them. That produced the 333 removals in #1031 and another 361 in #1035; re-run it when the population grows again, and verify by the gate, never by the classifier.

Three parsing facts the audit learned the hard way, for anyone writing tooling over these. An own-line phpcs:ignore covers its own line and the next one, not only the next. A // comment ends at ?>, so a sniff list must stop there (and at */) or the close tag is read as a sniff name — and a rewrite that eats it breaks every inline-HTML file. And ~33 annotations are the trailing form, on the same line as code ($result = $wpdb->get_row( // phpcs:ignore …): deleting the line deletes the statement, which is how one pass produced a parse error and 50 violations. Prose that merely mentions the token inside a docblock sentence is not a directive — only a token that opens its comment counts.

Comment-language guard (#1260)

tests/Unit/CommentLanguageTest.php fails when a comment line, or an i18n source string, carries two distinct Portuguese function words. It is the enforcement half of this file's opening rule, and it exists because the cleanup that rule triggered shipped incomplete three times in a row — each slice used an ad-hoc word list, and each failed differently: the first collided with English and code (era, com from .com, dos from DoS) and reported 959 false lines, so it was pruned; the pruned list was accented-only, so it never saw asercao / correcao / nao, which is how much of the test suite was written; and neither contained the 3-state tier names (só vê), so one slice left them and the next translated them. 35 files merged as "translated" while carrying a Portuguese clause mid-paragraph.

Two design decisions carry it, and both were measured rather than argued. Two distinct words, not one: a single word anywhere is what reported example.com and DoS, and the fix for such a false positive was always to delete the word — which is exactly how the recall went. Two words essentially never co-occur in an English technical sentence, so the list can stay wide. Quoted is mentioned, not used: a backticked or quoted span is stripped before matching, because a comment that OPERATES on Portuguese quotes it — all four exemptions the guard was first written with turned out to be that shape, so the rule replaced the allowlist instead of growing it, and ALLOWED ships empty.

Three things the tuning measured. The ground truth is the 160 lines #1278 removed, scored per BLOCK (the 60 contiguous comments they form), because half of those lines are fragments — * aquilo., // expectativa. — that no line-level rule can catch and none needs to: flagging one line makes the block findable and a human fixes the paragraph. Final: 59 of 60, zero false positives. A greedy tuner optimising recall alone proposed int and blank — English words, earning their place only by sitting on a Portuguese line in this corpus; that is the "dictionary describes itself" trap wearing a new costume, so every word is in for being Portuguese, never for fixing a block. And the wider list found six real Portuguese lines on a tree three passes had called clean, fixed in the same commit.

What it does not see, deliberately: a single Portuguese word — _x( 'ficha', … ) translating to itself is a product question (#1264), not a language defect — and the one block that is an English sentence with one Portuguese word used twice, which a two-distinct-word rule cannot reach by construction. Nor languages/ (that is the translation), CHANGELOG.md, commit messages or PR bodies; this file already records that the PR-body half is disciplined at the moment of writing, because nothing else catches it.

Catalogue-agreement guard (#1266)

tests/Unit/TranslationCatalogueAgreementTest.php compares the three files in languages/ with each other, by (msgctxt, msgid) and in both directions: every translated .po entry must appear with the same value in the .mo and the .l10n.php, and neither compiled catalogue may carry an entry the .po does not. Plurals compare form by form.

It exists because they had already disagreed in production, for two releases: the .po and .l10n.php carried #1209's fix while the shipped .mo still said Estado and Agradecimentos. WordPress reads .l10n.php only from 6.5 and this plugin's floor is 6.4, so on the floor the defect #1209 existed to fix was never delivered.

Parse the file with its own language where you can. The .l10n.php is read by including it — a side-effect-free return array() — which removes two traps outright, because PHP is the authority on how PHP escapes a string. The contrast with .github/scripts/ffc-uninstall-manifest.php, which parses uninstall.php as text, is deliberate: including that one would run the uninstaller. The rule is not "never include" — it is include when the file does nothing.

Two traps beyond the five the issue listed. The .mo keys a plural as msgid \0 msgid_plural while .l10n.php keys it by msgid alone, so a naive key reports every plural as drift. And a line-wise .po scan reports untranslated entries that are not — a multi-line literal's msgstr "" opens a continuation rather than declaring an empty string.

Source-coverage guard (#1284)

tests/Unit/TranslationSourceCoverageTest.php is the direction the one above cannot see. It compares the literal i18n calls in the source against the catalogue, both ways: A, every source string has a translated entry in the .po and the .pot, blocking at zero; B, every entry has a call site, a shrink-only ratchet whose register separates the structural (the plugin header is a comment, so no __() will ever emit Alex Meusburger) from the debt, each entry naming what replaced it.

It exists because languages/ was internally consistent throughout the two releases in which 79 on-screen strings had no entry at all and rendered in English on a pt_BR install (#1282). An entry that exists but is empty fails direction A too: msgfmt omits it, all three catalogues then agree it is absent, and pt_BR still renders the English source.

A mis-tuned extractor freezes a wrong baseline, and a baseline is believed. The first scan matched only T_STRING, so \__( 'I am not a robot', 'ffcertificate' ) — T_NAME_FULLY_QUALIFIED, one token carrying the backslash — was invisible: 30 strings, the whole ALTCHA block. Direction A was unaffected; direction B reported 52 orphans instead of 22, so the register would have declared 30 live strings dead. A canary pins that shape.

The extractor lives in tests/Support/I18nCalls.php and the .po reader in tests/Support/PoCatalogue.php, shared with the guard above for the reason .github/scripts/ffc-create-statements.php is shared. PoCatalogue::read() returns an ordered list, not a keyed map, because keying makes a duplicate msgid vanish rather than fail — and ffcertificate.pot carried one, which msgfmt would have rejected outright.

Schema-reachability guard (#1311)

tests/Unit/ActivatorSchemaGuardTest.php also fails when a table declared under includes/ has no declarer the Loader names. Activator::activate() runs on plugin ACTIVATION, and nothing else performs one — not an in-place WordPress update, not the rsync deploy to testes — so a module whose schema is reachable only from there gets its tables on a fresh install and on no upgraded one. That is not hypothetical: the two CSV-import tables of #1214 were in exactly that state, and the fresh-install job stayed green throughout, correctly, because it performs a real activation.

The rule is per TABLE, not per class, which is what lets it block at zero with no allowlist: a one-shot migration may legitimately stay unwired while the tables it declares are creatable through an activator that is wired. A class-level rule would have needed MigrationDynamicReregFields carved out by name.

Two things the measurement said that the issue did not. The exposure was five tables wide before anyone noticed, because reverting the fix makes the guard name seven — #1292 was simply the first commit to add a new table under a gap that already existed. And the alarm worked and nobody read it: the post-deploy smoke reported the missing tables on twelve consecutive deploys over about a day, which is the failure mode CLAUDE.md already warns about when it explains why a timeout stays green — an alarm people learn to skip is worse than none.

Healing the schema is not re-running activation, and UserDashboardActivator::maybe_migrate() is where that distinction is written down. Its create_tables() also creates the front-end dashboard page and registers a role; re-running those on every release would resurrect a page an administrator deleted and duplicate lifecycle the orchestrator owns. The method heals the two tables and nothing else. When wiring a new module, take the schema half only.

Deprecation-due guard (#1309)

tests/Unit/DeprecationDueTest.php fails when the plugin's own version reaches a recorded removal release, and fails when a WordPress deprecation call records none. It reads one marker, @removal X.Y.Z, written on the code that is still there — the convention, why the date cannot be prose, and the two token shapes the scan resolves are under "Deprecation cycles" in "Legacy and tech debt".

Its own self-check names a version, and that is the #1261 class inside a test. test_the_scan_finds_the_live_cycles() asserted assertContains( '6.27.0', $versions ) under a comment reading "Named rather than counted, so a shrinking cycle does not have to touch this file" — and closing the #1245 cycle in 6.27.0 invalidated exactly that line. Naming is not safer than counting here: both are claims about a value another file owns. The rule the fix wrote down is to name the oldest open cycle, because that is the one a reader most needs to see the scan still reaching, and to re-point it when it closes. A self-check cannot avoid naming something real — an empty scan must never read as clean — so this is maintenance the guard owes, not a defect to design away.

Config-level exclusions — audited verdicts

A rule switched off in phpcs.xml.dist, .stylelintrc.json, phpstan.neon.dist or phpunit.xml.dist is a phpcs:disable with repo-wide blast radius, so the same measurement was run over all four. The inline annotations are the audited part; these are recorded here so the verdicts are not re-litigated.

Generic.CodeAnalysis.UnusedFunctionParameter stays off — measured, not assumed. The verdict rests on the shape of what it reports, not on the total, and the total is not reproducible from a bare "enabling it": it depends on the standard the sniff is run under (vendor/bin/phpcs --standard=Generic --sniffs=Generic.CodeAnalysis.UnusedFunctionParameter includes/ reported 43 when this was written and 94 when #1261 re-ran it). Most of what it reports is structural: 14 are parameters consumed by a templates/*.php partial the method includes (the sniff cannot see across an include — and extracting markup to templates/ is this project's own convention, so the architecture creates this class), and 19 are signatures fixed by their caller (WP hook callbacks taking $atts / $hook / $update / $request, a guard interface's $ctx). Turning it on would mean 33 new annotations — exactly the churn #1035 removed. The 10 genuine ones were read and fixed in the same pass; re-run it as a one-off if the population grows, never as a gate.

Squiz.PHP.CommentedOutCode stays off. Every finding is a false positive and the verdict has survived a re-measure (9 findings when the rule was written, 16 when #1261 re-ran it over includes/, the same shapes throughout): the heuristic scores an explanatory comment as code when it quotes an enum (// 'daily' or 'span'), shows a shape (// [{day: 0-6, …}]) or names a token (// Remove {{ and }}). Zero real commented-out code.

Stylelint: declaration-block-no-shorthand-property-overrides and declaration-block-no-redundant-longhand-properties are ON (both measured 0 findings, so they cost nothing and guard a real bug class — a shorthand silently wiping an earlier longhand). no-descending-specificity and no-duplicate-selectors stay off; enabling both over the 28 non-minified sheets reports them in the hundred-and-few range for the first (103 when #1261 measured it) against a handful for the second, and every duplicate-selector hit is a deliberate section split — a token block and a layout block under the same selector, each with its own comment — where merging would destroy the intent.

The views/ carve-out is markup-only, and the two lists must stay in step. phpstan.neon.dist (excludePaths) and phpunit.xml.dist (<coverage><exclude>) both skip includes/admin/views and includes/settings/views — thousands of lines of markup carrying zero $wpdb, update_option, nonce or capability occurrences (grep the two directories for those five tokens; the answer is what justifies the carve-out, and it is still zero), i.e. the same rationale as templates/ above. includes/self-scheduling/views is deliberately NOT excluded — it holds real logic (capability gates, RequestInput reads, list-table wiring), so "a directory named views is markup" is not a repo-wide truth; judge by content. Two further paths, includes/views and includes/libraries, were listed for years and have never existed in the history — removed, since PHPStan's (?) optional-path suffix meant nothing ever reported them.

Test infrastructure

  • PHP: PHPUnit 9; tests under tests/Unit and tests/Integration.

  • JS: Vitest with jsdom (the pinned major lives in package.json, and Dependabot moves it — this file does not restate it); tests under tests/js/*.test.js. Real jQuery via jquery/factory bound to the jsdom window (tests/js/setup.js). Scripts under assets/js/ load via vm.runInThisContext so V8 coverage attribution survives (tests/js/helpers.js).

  • jsdom has no layout — jQuery :visible always reports false for shown elements. Assert on css('display') instead. Disable jQuery animation queueing in tests by setting window.$.fx.off = true in beforeEach so slideUp/slideDown/fadeOut apply immediately.

  • The pcov coverage-attribution gotcha is history, and the class_exists() preload it produced is now inert. During the #563 repository/god-class splits, pcov did not attribute coverage to a class first autoloaded during a test method, so a freshly-extracted class reported 0% while being fully exercised; the fix written down then was @covers \FQCN plus a class_exists( '\\FQCN' ) preload in setUp(), and 193 of the @covers files carried one when #1435 measured it, at 242 call sites — counting a bare class_exists( 'FQCN' ); STATEMENT, which is the part that has to be said. Grepping the call instead gives 205 files and 277 sites, because it also catches 31 if ( ! class_exists( … ) ) { … } guards that declare a stub when the class is absent (deleting one breaks the test) and AutoloaderTest's 3 assertions, where the call is the thing under test. The figure this sentence carried for three releases was 101 of 388, which no reading reproduces — this file's own opening rule failing inside this file. Re-measured and it no longer reproduces (pcov on PHP 8.4, PHPUnit 9.6.36, pcov.directory=./includes — the same ini coverage-shard sets, so this transfers to the gate): four preload-less files attribute normally (AccessRestrictionCheckerTest 82.1%, AdminUITest 96.0%, AppointmentEmailHandlerTest 91.6%, AppointmentValidatorBusinessHoursTest 30.0%), stripping the preload out of AdminClassTest and AdminSubmissionEditPageTest leaves them at 96.2% and 94.1% unchanged, and a canary built as the worst case — a test that names its covered class only inside the test body, no setUp at all — attributes 79.2%. Then re-measured over the whole population (#1435): all 242 preloads stripped, and each of the 193 classes run ALONE, twice, comparing the covered-statement count of its own isolated run — which is the reading that answers the question, because in a full-suite run another test covering the same class masks the effect entirely. Every class attributes the same number of covered statements with the preload and without it. One nuance is worth keeping, because it is the mechanism rather than a discrepancy: in CapabilityMigratorTest the denominator moves by 16 statements (57,255 against 57,271, stable across three runs each), so the preload does still change which files pcov instruments — it simply no longer changes what pcov ATTRIBUTES. And one of the 193, TabReregistrationTest, errors under --filter in both states for a reason that predates this (the order-dependence this section already describes), so it is evidence for neither side. So do not add a preload to a new test, and do not sweep the existing ones out either: they cost a line, they are the fallback if the driver regresses, and deleting them is churn against a property that now holds without them. Attribution is still spot-checkable the same way: php -d pcov.enabled=1 -d pcov.directory=./includes vendor/bin/phpunit --coverage-clover /tmp/cov.xml --filter <Test>.

  • A function_exists() guard makes tests order-dependent, and the symptom is remote. LabelSorter::locale() guards on function_exists( 'get_locale' ), so whether it takes the WordPress branch or its en_US fallback depends on whether any earlier test in the process stubbed that function — Brain\Monkey defines it via Patchwork, and it stays defined afterwards. Adding a get_locale stub to one test class therefore breaks unrelated classes that run later, with "get_locale" is not defined nor mocked in this test pointing at a file the PR never touched. Both times this happened (#1053), the fix was to stub the function explicitly in every test that reaches the guarded code, so the branch is chosen rather than inherited. Defining it in tests/bootstrap.php does not work — Patchwork cannot redefine a function declared in an uninstrumented file, so every Functions\when() for it starts throwing instead. When you add a stub for a WP function that some function_exists guard checks, expect the blast radius to be "every test after this one, alphabetically".

  • Stub the minimum the code under test actually resolves — every function you teach Patchwork stays taught for the rest of the process. The bullet above covers a stub that is missing; this is the same machinery seen from the other side, and it cost three separate rounds in one session (#1236). A stub is not scoped to the test that declares it: Functions\when() makes Patchwork define that function, and after teardown the definition remains without an expectation, so any later test reaching it fails with "…" is not defined nor mocked in this test — pointing at a file the PR never touched. Three shapes, each of which looked like a different bug:

    • One too many, global. A dbDelta stub added for a new activator test broke ActivityLogTest, which had relied on that function never being defined. The fix was not to stub it everywhere: it was to stop needing it — the $wpdb double was taught to answer SHOW TABLES LIKE with the table name, so no chain reached a CREATE at all. That also made the test measure the case that matters, an established install.
    • One too few. Gating four activator chains on an option (#1231) made them call get_option/update_option, which the existing activator tests did not stub — 16 errors, all legitimate, all fixed by stubbing in setUp with a value that deliberately does not equal FFC_VERSION, since those tests exercise the body.
    • One too many, namespaced — the subtle one. Stubbing Ns\get_option creates that namespaced function, and from then on an unqualified get_option() inside Ns stops falling back to the global one. A test that stubbed only the global (RewriteHtmlImageRefsMigrationStrategyTest) broke, with no relation to the change. The namespaced stubs were never needed — PHP's fallback already resolves the global — and copying them "for symmetry" with a sibling test is how they get reintroduced, so the reason lives in that test's setUp.

    The corollary is what costs the time: none of the three reproduces in the file alone. All three passed under --filter on the affected file and only appeared in the full suite. --filter is not evidence here; the run that decides is vendor/bin/phpunit with no arguments. And when a failure does appear, the decisive experiment is cheap — run the suspect test together with the victim, then the sibling that does the same thing together with the victim: if the sibling's pair passes and yours fails, the difference is in your stubs, not in the victim.

  • A closure parameter taken by reference breaks under Patchwork, and only sometimes. Patchwork instruments every dynamic call and dispatches it through call_user_func_array(), which cannot pass by reference — so $add = static function ( array &$bucket, … ) dies with Argument #1 ($bucket) must be passed by reference, value given at CallRerouting.php. The dispatch only happens once some patch is live in the process, so the file passes alone and fails after any Brain\Monkey test that ran before it: a --filter scoped to names alphabetically after the file under test reports green while CI reports ten errors (#1177). Capture the array with use ( &$bucket ) instead — a captured variable is not a parameter, so the dispatch never touches it. And when reproducing an order-dependent failure, pick a filter that keeps a patching class before the one you are testing.

  • templates/ is outside the coverage scope. phpunit.xml includes only ./includes for coverage, so extracting inline markup from a god-class into templates/*.php partials reduces the class without touching the coverage floor (the F1/F2 lesson). Logic stays in includes/ (and stays covered); pure markup moves to templates/.

  • tests/ is tab-indented, and that is enforced. phpcs.xml.dist excludes tests/ on purpose — WordPress-Docs would demand a file, class and method docblock in every file under it. Indentation is the one exception: phpcs-tests.xml.dist carries exactly two sniffs (DisallowSpaceIndent, ScopeIndent) and runs over the whole directory in the PHPCS workflow. The suite was 250 space-indented files against 170 tab-indented ones with nothing enforcing either, so a one-off reformat would have drifted back; the gate is what makes the reindent worth doing. Do not widen that ruleset — line length, docblocks, naming and Yoda conditions are how it stops being cheap. Two mechanics to know. The ruleset must set <arg name="tab-width" value="4"/>: without it PHPCS measures a tab as one column and reports every tab-indented line as expected 1 tabs, found 1 spaces — 105,985 errors, the gate failing the files it exists to approve. The main ruleset never needed it explicitly because WPCS sets it inside WordPress-Core; this one references only Generic sniffs. And the reindent commit is listed in .git-blame-ignore-revs, so git blame still names whoever wrote a test rather than the reformat — GitHub honours that file automatically, locally it is git config blame.ignoreRevsFile .git-blame-ignore-revs once per clone. Add a commit there only when it is genuinely whitespace-only.

  • Running the suite locally. Install pcov once (apt-get install -y --no-install-recommends php8.4-pcov); the full suite is ~10 min. The full run is required evidence only when the change can reach another file — a stub added or removed (Functions\when()), a shared tests/Support/ helper, a renamed identifier, or anything in includes/. That is what the order-dependence rule above is scoped to: --filter is not evidence there. For a change that provably cannot pollute the process — a comment-only or assertion-message-only edit, where the token skeleton is byte-identical — the full local run buys nothing the --filter run and CI do not already give, and costs ~9 minutes per slice. Measured over the #1260 slices: two of four were comment-only and the full run found nothing in either. Do not skip it on the strength of intent — skip it on the strength of the skeleton diff being empty. Scope while iterating with vendor/bin/phpunit --filter <Test>; vendor/bin/phpstan analyse --no-progress <path> and vendor/bin/phpcs --standard=phpcs.xml.dist -q <files> (auto-fix with vendor/bin/phpcbf) reproduce the PHPStan/WPCS gates.

  • PHPStan/phpdoc idioms. Put @phpstan-type on the class docblock (not the file docblock); consumers use @phpstan-import-type X from Y on their own class docblock. Avoid @todo ( and other @tag ( openings — phpdoc parses the ( and errors; use prose instead.

Build & assets

When editing assets that ship to the browser, rebuild the minified bundles so the matching *.min.* and .map files stay in sync — the Verify minified assets are up to date CI job fails otherwise:

  • JS → npm run build:js (regenerates assets/js/*.min.js + .map).
  • CSS → npm run build:css (regenerates assets/css/*.min.css + .map).
  • Both at once → npm run build.

3. Architecture and patterns

Repository pattern (Reader/Writer split)

Data-access classes are split read-side vs write-side. Two shapes — pick by whether the repo needs multi-statement transactions:

  • Static repos (the common case). Reads live in *Reader, writes in *Writer; both use \FreeFormCertificate\Core\StaticRepositoryTrait and return the same cache_group(), so a write invalidates the caches a read populated. Callers call *Reader:: / *Writer:: directly — there is no façade. The nine static façades (CustomField, Audience, AudienceBooking, RecruitmentCall/Candidate/Notice/Adjutancy/Reason, ReregistrationSubmission) were retired in #563 B3-A (#594). The row-shape @phpstan-type and any public constants (field-type lists, status sets, default colors, …) live on the Reader as the canonical home; consumers do @phpstan-import-type … from …Reader.
  • Instance repos. AppointmentRepository, SubmissionRepository, UrlShortenerRepository are kept as thin façades that extend AbstractRepository, compose a *Reader + *Writer in the constructor, and expose the inherited generic CRUD. They are the transactional aggregate root: begin_transaction() → FOR UPDATE read → write → commit() run on the one shared global $wpdb the façade/reader/writer all bind. Do NOT retire these into separate call sites — that coherence is the whole point (deliberate B3-A decision; reiterated in each class's docblock).

Legacy third shape — single-class instance repos (not a template for new code). Four older classes extend AbstractRepository directly for the inherited generic CRUD + cache helpers (implementing only get_table_name() / get_cache_group()), with no *Reader/*Writer split and no transactions: CalendarRepository, BlockedDateRepository (ffc_self_scheduling_*), FormRepository (wp_posts metadata), UserProfileRepository (ffc_user_profiles, extracted from UserManager/UserCreator in #340). They are neither the static pattern nor transactional aggregate roots — they predate the static-default guidance and are left as-is (converting them to static would be pure namespace churn against the module-boundary baseline for zero functional gain). So the extends AbstractRepository census is 13 classes = these 4 + the 3 transactional façades + their 6 composed *Reader/*Writer halves. Don't cite these four as examples of the instance-façade pattern, and don't add new ones — new data-access code still follows the static default below.

When splitting or adding a repo, default to static; reach for the instance-façade only when callers need a read-modify-write transaction. Test note: a caller's alias mock that stubs both reads and writes must become two alias mocks — one on the *Reader, one on the *Writer — or the write call hits the real (un-mocked) class.

Module bootstrap (per-module loaders)

Each feature module exposes a single bootstrap entry point — a *Loader class whose init() wires the module's runtime classes — so the orchestrator (Loader::init_plugin()) touches one symbol per module instead of newing-up its internals inline. This keeps Loader a thin composition root and narrows the Root → <module> dependency surface (#563 B3 coupling reduction). Current loaders: AdminLoader, AudienceLoader, RecruitmentLoader, UrlShortenerLoader, ReregistrationLoader, SelfSchedulingLoader.

Pattern: init() runs the module's wiring in its original order, gating admin-only pieces behind is_admin(); held-alive instances are kept as protected properties (so PHPStan doesn't flag them write-only); fire-and-forget ::init() / one-shot new calls need no property. Pin the wiring with a *LoaderTest (overload/alias-mock the wired classes; assert each is constructed / ::init()-ed) carrying @covers. The existing loader tests also carry a class_exists() pcov-preload in setUp(); a new one does not need it — see "The pcov coverage-attribution gotcha is history" in "Quality gates and testing".

Only extract genuine module bootstrap. A loader is worth it when the Root → module edge is bootstrap wiring (instantiating the module's runtime classes). It is NOT worth it when the edge is orchestrator-level lifecycle — role/capability registration & migration (RoleRegistrar / CapabilityManager / CapabilityMigrator), cron-event registration, or activation/upgrade maybe_migrate() — which legitimately belongs to the orchestrator. Extracting those into a "loader" would be indirection that narrows nothing (the bad-facade trap; see "Repository pattern" for the same principle).

Documented exception — UserDashboard has no loader (deliberate, B3 phase 2). Its only bootstrap wiring is AccessControl::init() + UserCleanup::init() (2 calls, left inline in Loader). The bulk of its Root → UserDashboard surface is capability/role lifecycle (RoleRegistrar / CapabilityManager / CapabilityMigrator) invoked from Loader::register_ffc_roles_safe() / ensure_*_caps() — orchestrator responsibility, not module bootstrap. A UserDashboardLoader would move 2 lines without shrinking the edge, so it was intentionally skipped.

Shared-service module directories

A few includes/ modules are small, single-purpose service buckets whose names are generic enough to invite drift. Keep each scoped to its stated purpose — do NOT let it become a "misc" drawer, since renaming to a crisper namespace later would ripple through the module-boundary baseline, every use / @covers / alias-mock and the autoloader for only a cosmetic gain (the same namespace-churn cost as any cross-module move). The guard is scope discipline, not a rename:

  • services/ (\Services) — user-centric query/identity services only (UserService, UserIdentifiersQueryService). A service that isn't about users belongs in its own domain module, not here.
  • integrations/ (\Integrations) — adapters to external systems (EmailHandler → SMTP, IpGeolocation → geolocation API). A class with no outbound/external dependency is not an integration.
  • scheduling/ (\Scheduling) — cross-cutting scheduling-domain services shared by the self-scheduling and audience features (DateBlockingService, WorkingHoursService, IcsGenerator, SchedulingMailer). Distinct from the self-scheduling/ and audience/ feature modules and from the "Scheduling" admin menu: feature UI/handlers go in those modules; only shared scheduling logic lives here.

Email architecture (one pipeline — #662)

Every plugin-composed email flows through one shared pipeline; never hand-roll a chrome, a transport, or an inline body. When adding or touching an email:

  1. Compose the email body only (the inner content). Resolve placeholders with Core\TokenResolver::resolve() ({{token}}) — never hand-rolled str_replace.
  2. Wrap in the single configurable chrome: EmailHelperTrait::ffc_email_document( $body, array( 'recipient' => $to ) ). The chrome (header/body/footer/wrapper) is admin-configurable via Core\EmailTemplateOptions (the "Email Model" box, Settings → SMTP) and rendered by templates/emails/layout.php (table-based, inline styles for Gmail/Outlook). There is exactly one chrome — SchedulingMailer::wrap_html and the per-email cards were retired.
  3. Send through the chokepoint Core\EmailService::send() (or EmailHelperTrait::ffc_send_mail(), which sets text/html). Never call wp_mail() directly. The global "disable all emails" kill-switch is enforced inside EmailService::send() — do not re-gate it caller-side.
  4. Default bodies live in files, one per case, under templates/emails/. Return-array files (return array( 'body' => __( … ) ), loaded via the allowlisted Core\EmailTemplates::load() / ::body()) for token/editable-default bodies; echo partials (via ffc_render_email_partial()) for handler-built ones. Never build email HTML inline in a handler class — extract it to a templates/emails/*.php file.
  5. Editable emails edit the body only, never the chrome. Use wp_editor (TinyMCE) + a "Restore Default Text" button wired by the generic assets/js/ffc-email-restore-default.js (data-editor + data-default-key; default supplied via wp_localize_script['ffcEmailRestoreDefaults']).
  6. Surface the P5 notice: Core\EmailDisabledNotice::render() at the top of any admin surface that edits an email.

When consolidating or auditing, grep every send site (EmailService::send / ffc_send_mail / wp_mail() and confirm none bypasses ffc_email_document — the #662 audit found emails the roadmap table had missed (a plain-text calendar-deletion cancellation, two admin notifications). WordPress-core emails (wp_new_user_notification, password resets) are out of scope — they use WP templates, not our chrome.

CSV export architecture (one contract, two adapters — #772)

Every plugin CSV export flows through one source contract with two delivery adapters; never hand-roll header() / fopen('php://output') / exit, an AJAX job loop, or a temp-file scheme. A source owns only its domain specifics (auth, columns, row formatting, the query); the Core adapters own the delivery lifecycle.

The contract (Core). A source implements exactly one of two interfaces, both in includes/core/:

  • SyncSourceInterface — bounded, provably-small output. Methods: authorize() · filename() · header() · rows(): iterable. Delivered by Core\SyncCsvExport → Core\CsvStreamer → Core\HttpCsvDownload (streams straight to the browser in one request; CsvStreamer is the synchronous-delivery adapter — pure orchestration over the injectable CsvDownloadInterface, so it is unit-testable with a buffered double). Use for outputs whose size is bounded by construction (audit-log ring buffer, static example/sample templates, the audience import/export whose dedup+hierarchy don't survive a cursor) or bounded by a cap the source itself imposes, when synchronous delivery is what buys the property that matters (#1295: the identity audit's findings name accounts, and the temp file is what the batched path cannot avoid). That third case is not the default — the rule below still says batched when in doubt — and it carries an obligation the first two do not: a cap that is reached must be REPORTED in the output, or a partial list reads as a complete one.
  • BatchedExportSourceInterface — timeout-safe for datasets of unknown/large size on shared hosts (where set_time_limit() is blocked). 16 methods: per-phase authorize_start/authorize_batch/authorize_download, job_owner_fields, sanitize_filters, count, build_context, header, filename, fetch_page(cursor, size), cursor_of, format_row, extra_start_response, on_complete, on_before_download. Delivered by Core\BatchedCsvExport (start → batch×N → download over a temp file + transient job). Use for every data export (submissions, public forms, url-shortener, activity-log, audience-bookings, appointments, reregistration).

Registry + dispatcher (dependency inversion). Core owns the interfaces, the two engines, Core\SourceRegistry (type → lazy factory) and Core\BatchedExportDispatcher (one AJAX trio ffc_export_start / ffc_export_batch / ffc_export_download, priv and nopriv, routing by the type field to the registered source, which owns its own authorize_*). The concrete sources live in the feature modules and register a factory at bootstrap (SourceRegistry::register('audience_bookings', fn() => new AudienceBookingExportSource())). The engine/dispatcher never reference a feature class → no Core → feature edge (the → Core edges already exist), so ModuleBoundaryTest stays green without regenerating the baseline.

When adding or touching a CSV export:

  1. Pick the adapter by cardinality: bounded/static → SyncSourceInterface; unknown-size data → BatchedExportSourceInterface. When in doubt for a data export, batched.
  2. Register the source from the module's bootstrap under is_admin() (true on admin-ajax, so the source is reachable during the export job) — in the module *Loader, or a class the module always constructs on admin requests. Never register from an admin_menu-hooked method: admin_menu does not fire on admin-ajax.php, so the source would be missing exactly when the dispatcher needs it.
  3. Keyset, never LIMIT/OFFSET, for batched fetch_page: WHERE … AND id < %d ORDER BY id DESC LIMIT %d, with cursor_of() returning (int) $row['id'] and the first page cursor PHP_INT_MAX. The id desc order is stable across concurrent inserts during a long export — this is a deliberate behaviour change from any prior on-screen orderby (call it out in the CHANGELOG). Add a count/keyset method to the domain's Reader (not the source); when a shared count() can't honour a filter (e.g. a JOIN-only column), add a dedicated count_for_export rather than reusing it.
  4. Freeze dynamic columns in build_context. When columns are per-row JSON (submissions, appointments) or per-campaign fields (reregistration), resolve the full column set once at start and carry it in the frozen job context — the header must not drift between batches. For row-JSON, scan keys with a lightweight keyset pass over only the JSON columns.
  5. Front-end: drive the button through the shared JS. The "Export CSV" button calls window.FFCBatchedExport.run({ type, ajaxUrl, nonce, startData, callbacks }) from assets/js/ffc-batched-export.js — never a bespoke start/batch loop. Enqueue ffc-batched-export (dep ffc-core) on the page and localize the job nonce. Register the click handler at a point that always runs — its own method / top-level delegated $(document).on('click', '#…', …), not folded into another init*() that early-returns on that page (that silently dropped the audience-bookings handler — #783). startData may be an object (admin) or a $form.serialize() string (public); the driver appends &type= to a string and object-merges otherwise — don't hand it a string through $.extend (that spread its chars and dropped the page nonce — #781).
  6. Ownership is pluggable via authorize_*: capability + job nonce + user_id fence for admin sources; UUID + per-job nonce + IP-hash for the anonymous public source. Admin sources gate a dedicated ffc_export_* cap via Capabilities::current_user_can_admin_or().

PII-on-disk is accepted for the batched sources that decrypt (submissions, appointments, reregistration, public): the temp file lives under wp_upload_dir()/ffc-tmp (writable, survives plugin updates, per-site on multisite), guarded by a .htaccess Deny from all, random-UUID filenames, unlink after download, and a daily cleanup cron. An export stays synchronous when that is what keeps its rows off disk (eliminating the risk rather than mitigating it) — the audit log, and the identity audit's findings (#1295). The second is the one that generalises the reasoning, because it had to argue against the default: its population is not provably small, only capped, so the choice was between a temp file holding which accounts share an identifier and a cap that says when it was reached. An install that hits the cap is the trigger to move it to the batched engine — not to raise the number, and not to build that engine on spec (the #788 / #902 / #993 criterion). Nginx caveat: .htaccess is Apache-only — on nginx the temp dir must be denied at the server block (e.g. location ^~ /wp-content/uploads/ffc-tmp/ { deny all; }, or keep wp-content/uploads non-executable and unreadable for that path). Document this for any nginx deploy; the in-code .htaccess does not cover it.

When auditing, grep every legacy output site (php://output, fopen(, header( 'Content-Type: text/csv, CsvStreamer, ->stream() and confirm each export routes through a source + one of the two adapters — the #772 consolidation retired seven bespoke exporters this way. One deliberate exception survives the grep: Frontend\PublicCsvExporter::stream_form_csv() still writes to php://output directly. It is the graceful-degradation half of the only public/frontend export — a plain <form> POST to admin-post.php that works with JS disabled (admin exports need no such fallback; they run in wp-admin where JS is assured). It is not a SyncSourceInterface because it is a direct admin_post streaming handler (not an AJAX job) and it emits an HTML 413 page over the row cap, which does not fit the rows(): iterable contract. Leave it bespoke — do not "fix" it to satisfy the grep.

The no-JS half is now conditional, and that is deliberate (#1053). Under the ALTCHA-only captcha mode this path exists but cannot be completed: the widget needs JavaScript, so a visitor without it has no way to pass the challenge and handle_request() refuses. That is an accepted consequence rather than a regression — the audience for this export is operators working on a desktop, so the mode is defensible for them — but it does mean "works with JS disabled" is a property of the math and composite modes only. Keep the handler: it is the whole no-JS path in the two modes that have one, and it is also the rollback that needs no redeploy.

Batched migrations — what makes a batch stop (#1378)

A migration card runs its batch again and again until the batch reports it did nothing. So termination is a property of what execute() counts, and counting the wrong thing produces a loop that no test sees and no gate catches, because everything involved is individually correct.

Measured, on the testes site. The certificate-capability card reported TOTAL 10 · PENDING 1 while its button had processed 1,848 records — about 185 passes over the same ten accounts. execute() returned the number of accounts selected, not the number changed; CapabilityManager::grant_certificate_capabilities() skips a capability the user already holds, so a grant that did nothing still reported progress. The method's own docblock asserted the opposite — "a batch that grants nothing therefore means the work is done" — and the code could never produce that condition.

The account it could never clear held the capability through a role. A role's capabilities live in the wp_user_roles option; a user's wp_capabilities meta carries the role NAME plus per-user grants and nothing else. So a LIKE over that meta cannot see the capability, while has_cap() resolves it fine: the pending query and the grant were asking different questions, and an account could sit in the pending set that the work would never move. It stayed invisible through the 1,478-account repair because ffc_end_user carries no FFC capability at all — there the grant really is per user, and the LIKE is the right question.

So the rule is about the counter, not about the query. A cursor-free batch terminates only when its processing provably negates its own pending predicate — and proving that for every future edit is not something a reviewer can keep doing. Count what changed: then a batch that changes nothing returns zero and stops, whatever the predicate later drifts into. Termination must not depend on the measure being right, because the measure is the part that goes wrong.

All eleven strategies were checked against it, and each of the other ten is safe by one of three means — worth knowing which, because they are not interchangeable:

  • A cursor or keyset (identity_normalization, both key rotations, email_hash_rehash, identity_index_backfill, activity_log_clear_plaintext) — advances whatever the work does.
  • A write that negates the predicate (cpf_rf_split) — it filters on cpf_rf_hash IS NOT NULL and writes NULL into that very column, and increments only on a successful update(). Both halves are needed: the second is what the looping card lacked.
  • An explicit visited register (rewrite_html_image_refs, import_legacy_templates) — the key is recorded even when the item fails, so progress does not depend on success.

The looping card had none of the three.

The neighbouring rule, from the same arc: a scan that read nothing must never render as clean. IdentityConflictQuery::rf_check_digit_failures() returns failures only, and returns none when no store carries the three columns it needs and when nothing decrypts — which is the ordinary state of a staging site whose key does not match the data. A screen that prints "every stored RF satisfies its check digit" off an empty list is claiming something it never measured. That is the #1071 / #1094 rule reaching past the guards it was written for: it binds any surface that reports an absence, not only CI.

Captcha architecture (one contract, three modes — #1053)

Every public form is guarded by the same block: a honeypot (provider-independent, owned by Core\SecurityService) plus a challenge from whichever strategy is configured. Never verify or render a challenge directly — go through the contract.

The contract is Core\Captcha\CaptchaProviderInterface: id() · render_fields() · verify() · peek() · challenge_payload(). Core\Captcha\CaptchaProvider::resolve() picks the strategy from ffc_settings['captcha_provider'] and falls back to math on an unknown value — a typo in an option must not take the public forms down. Three strategies: MathCaptcha (math), AltchaCaptcha (altcha), CompositeCaptcha (both, ALTCHA with the math half inside <noscript>).

verify() spends the challenge; peek() checks it without spending. This distinction is the contract's, not one strategy's, and it exists because of a live bug: the public CSV download is the plugin's only two-request flow and validated the same token twice, so single-use tokens made the second request reject the answer the visitor had just been told was correct (#1061). The rule: the challenge is consumed by the action it authorises, not by the metadata read that precedes it. A peek() that merely skips the ledger is wrong — it must refuse an already-spent proof too, or the contradiction moves one request downstream. Any new strategy implements both.

Composite mode is accessibility, not security. The server accepts either proof, so an attacker picks the cheaper one and the effective strength equals the math challenge's. Say so wherever it is offered; do not let it be read as "both, therefore stronger".

Five sites issue a challenge, and all five go through the provider: the render sites (SecurityService::render_security_fields()), the four retry paths (SecurityService::with_fresh_challenge()), and the cached-page fragment refresh (Frontend\DynamicFragments). The last one was missed once and is the reason this is written down — its client half dispatches on the payload's provider key and leaves an unrecognised one alone rather than half-applying it.

ALTCHA specifics that are not guessable and cost time to rediscover (all verified against the vendored 3.2.2 bundle, not documentation):

  • The element accepts exactly nine attributes — auto, challenge, configuration, display, language, name, theme, type, workers. Everything else (hideLogo, hideFooter, humanInteractionSignature, setCookie, the floating options) travels as JSON in configuration. Written as an attribute it is ignored in silence — no console error, no visible failure. Core\Captcha\CaptchaSettings is the one place that knows which is which.
  • Translations are not an attribute either. They live in the globalThis.$altcha.i18n store, keyed by language and selected by language. assets/js/ffc-captcha.js registers the plugin's own strings there, which is what keeps the upstream 52 KB i18n bundle out of the page.
  • maxnumber does not exist in 3.x — the word is absent from the bundle. The solver counts up without a ceiling until it reproduces the hash, so difficulty is the size of the server's secret number and nothing else (expected work ≈ half of it). It is still emitted in the challenge because the published format carries it and other clients read it.
  • The widget refuses to run outside a secure context (isSecureContext), throwing rather than degrading. The ALTCHA-only mode is therefore blocked at save time without HTTPS; the composite mode is allowed because its <noscript> half still works.
  • Vendor the UMD build (libs/js/altcha-<version>.umd.js). The ESM build needs <script type="module">, which only wp_enqueue_script_module() emits — WP 6.5+, above this plugin's declared floor of 6.4.
  • The wire format is ALTCHA's original ("v1") challenge. Given a top-level challenge key the 3.x widget marks it _version: 1 and posts back {algorithm, challenge, number, salt, signature, took} in base64, with the counter hashed as a decimal string. The expiry rides in the salt as ?expires=<unix>; the salt is not signed directly and does not need to be, because challenge is the hash of salt plus secret.

Keys, bounds and privacy defaults live in CaptchaSettings — allowed values for each attribute, and the four bounded numbers. Each is clamped on read as well as on save, so a value written before a bound moved still lands somewhere the widget can cope with. The last two are the challenge-issuing cap and its window (#1111) — a cap whose window is a private constant is half a setting, since the same number means something different over a minute and an hour — and they are the one exception to "a captcha setting lives in ffc_settings": they are stored in ffc_rate_limit_settings under ip.captcha_max_per_window / ip.captcha_window_seconds and edited on the Rate Limit tab, because that tab groups by what is limited and this is limited per IP — next to the submission cap an administrator raises for the same institutional-NAT reason. Bounds stay in CaptchaSettings; storage belongs to the tab that owns the option. 0 on the cap means no cap, the same convention the per-endpoint read limits use, and it is the only honest answer for heavy NAT, where any per-address number caps the building rather than the farmer — 0 on the window is not a value and floors like any other out-of-range number. The window's floor is 1 second, not 60: under a minute the throttle is mostly decorative (a caller just paces across boundaries), but that is a recommendation the field states, not a range the code refuses — the bounds here exist against a value the runtime cannot use, never to express a security opinion the administrator is not allowed to overrule. Lowering it required fixing what a cleared field writes: (int) '' is 0, so every bounded numeric setting in the admin silently stored its own floor when emptied, and with a floor of 1 that would have meant a one-second window. Both write paths now treat empty as "not supplied" — the tab's rebuild-save keeps the stored value, and SettingsAjaxEndpoint refuses the write with a 400 (for every int key, not just these) so the field's badge shows the failure instead of a value nobody typed. The window length is part of the transient's bucket key, so changing it moves callers to a fresh bucket instead of reinterpreting counts taken under the old one. humanInteractionSignature (pointer and keyboard timings) and setCookie are forced off and deliberately not configurable: these are public-sector forms under the LGPD, the proof of work already carries the anti-automation load, and an administrator toggling them back on would change what the site has to disclose without being told so.

Signing and single use are shared, not per-strategy: Core\Captcha\ChallengeSigner (key derived from wp_salt('nonce') — no option, nothing for uninstall.php or the fresh-install manifest) and Core\Captcha\ChallengeStore (transient ledger; redeem() spends, is_spent() reads). A new strategy reuses both rather than inventing its own.


4. Stylesheets and theme

The plugin is a guest in someone else's document — the site's theme owns <html> on the frontend, WordPress owns it in wp-admin — and almost everything below follows from that one fact. It is its own section because "Architecture and patterns" is about how the PHP is composed and this is about what reaches the browser; a reader answering a question about repositories, email or CSV does not need the stylesheets, and a reader answering one about a selector should not have to find it inside a section named for something else.

Four subsections, in the order a question usually arrives: how many sheets there are and why, how a class is named, how a rule is scoped to a screen, and how colour works. Each records the measurement behind a standing decision, so a proposal framed as "adopt BEM" or "collapse the sheets into a bundle" is answered from what was measured rather than from scratch.

Stylesheet architecture (28 sheets, one per screen — deliberately not a bundle)

assets/css/ holds 28 non-minified sheets, 432 KB — 17 admin-only, 7 frontend-only, 4 both — and each is enqueued by the screens that need it. wp_enqueue_style is the entry-point mechanism. There is no bundler and no preprocessor: npm run build:css runs cleancss over each sheet in place, and build:js does the same with terser. The stack is plain CSS, PHP templates and jQuery.

Per-screen, not two bundles — measured, not preferred. Collapsing the sheets into admin.css + front.css (the usual advice for an app you own) costs roughly double on every screen:

Screen Today Single bundle
Appointments (admin) 113 KB · 7 sheets 220 KB (1.9×)
User dashboard (frontend) 99 KB · 4 sheets 218 KB (2.2×)

The bytes are the smaller reason. Conditional loading is what keeps the plugin's CSS off the screens that are not ours — a single admin bundle would land on users.php, the post editor, and any screen whose gate happens to match, which is precisely the blast radius CssNamespaceAnchorTest exists to bound. Granularity here is a safety property, not an optimisation.

The plugin is a guest in someone else's document, and that settles two whole layers. On the frontend the surrounding document belongs to the site's theme; in wp-admin it belongs to WordPress. So:

  • No reset, ever — no normalize, no global box-sizing, no * { margin: 0 }. Measured: zero in the codebase, and it must stay zero. A reset is for whoever owns <html>.
  • No bare element selectors on the frontend — p, a, h1-h6, body styled unqualified would repaint the host theme. The invariant is stronger than the frontend now, and it is re-measurable: across all 28 sheets, no selector is fully unqualified — every one carries a class, an id or an attribute somewhere (run CssSelectors::of() over the sheets and reject anything matching no [.#[]; :root is the single allowlisted exception, for the reason the anchor ratchet gives). Five admin-side selectors once were unqualified and sat in #1152's baseline; #1184 and #1202 anchored them rather than deleting them — code became .ffc-admin-page code, pre code became .ffc-admin-page pre code, .form-table th became .ffc-settings-wrap .form-table th — which is what emptied that baseline. Reintroducing one is architecture to argue for, not debt to inherit.

How sheets relate. ffc-common.css is the root: it declares the whole --ffc-* palette on :root, the wp-admin base pair, and the shared components. The line and selector counts that used to sit in this sentence are gone on purpose — they drifted with every edit and carried nothing the sentence needed. Every module sheet declares it as a wp_enqueue_style dependency, which is both what makes a module's override of it deterministic and what AdminStylesheetTokensTest direction B enforces. The admin base is a chain — ffc-pdf-core → ffc-common → ffc-admin-utilities → ffc-admin-css → ffc-admin-submissions-css — and a dependency edge is the only thing that fixes the order between two sheets; without one it falls to enqueue order, i.e. the order modules are wired in Loader (see the component-ownership guard in "Quality gates and testing").

Three facts about that graph that cost time to rediscover. Two handles can name one file: ffc-admin.css is ffc-admin on users.php and ffc-admin-css on FFC screens, with different dependency lists. One handle can be enqueued with different dependency lists at different sites (ffc-admin-settings), and WordPress keeps whichever registered first. And a gate can be wider than its handle's name: ffc-admin-submissions-css is gated on is_ffc_page(), which matches any ?page=ffc-* screen, so it loads far beyond the submissions page its docblock names.

Four guards hold this shape, and each is described in "Quality gates and testing": CssNamespaceAnchorTest (every selector names something we own), StylesheetOwnershipTest (two sheets declaring one class must have an edge), AdminStylesheetTokensTest (a sheet reading tokens declares ffc-common; no colour literals), DarkModeCssTest + TypographyTokensTest (the palette and the scale). The CSS parser they share lives in tests/Support/CssSelectors.php.

Standing decisions, so an ITCSS-shaped proposal is not re-evaluated from scratch. The layered advice (settings · tools · generic · elements · objects · components · utilities, with per-area bundles) is sound for a standalone app with a bundler and a document of its own. Mapped onto this plugin it splits three ways, and the split was measured rather than argued:

  • Already in place, by another route. Settings — the token palette and the typography scale, as native custom properties, which beat preprocessor variables here because they respond to the dark theme at runtime. Utilities — ffc-admin-utilities.css. Entry points — per screen, which is finer than per area.
  • Would be a defect here. Generic (a reset) and Elements (bare tag styling) both assume ownership of the document, which a plugin does not have; the second is the very thing #1152 froze as debt. BEM is churn: the ffc- prefix already carries the isolation BEM's naming would, over ~1,000 classes. What is worth keeping from that advice is component-specific naming, which #1151 / #1154 / #1162 have been doing one component at a time.
  • Genuinely worth doing — and done as of #1184. A page-scope class on the admin wrap, which is the anchor #1152's baselined selectors need in order to be fixed. See "Page scope on the admin wrap" below.

A preprocessor has one argument that sounds real — a mixin for the same visual shape under different owners — and the example it was stated over does not survive measurement. The .ffc-status-badge declared bare in three admin sheets looked like one shape three modules were forced to copy. It is not: 3px 8px / 2xs / uppercase, 2px 8px / xs / uppercase + line-height, and 2px 8px / xs / no uppercase + nowrap. Three different badges wearing one name — which is a rename in the #1151 / #1154 / #1162 mould, not a mixin. #1168's shared skin object does not resolve those two baseline pairs either, because what collides there is the structure, not the colour pair.

The duplication that is real is the skin pair itself — background: var(--ffc-X-bg); color: var(--ffc-X-text) and nothing else in the rule. #1168 counted 57 such rules across 13 sheets, plus 37 more carrying the pair inside a larger component; the renames of #1170 / #1176 have moved those figures since, so re-count before quoting them (the reading is: a rule whose only two declarations are a --ffc-*-bg background and a --ffc-*-text colour). What has not changed is the shape, and it has a pure-CSS answer (#1168) that needs no build step, so the preprocessor buys nothing here. Revisit only if a shape appears that a shared class cannot carry, and decide it the #788 / #902 / #993 way — from proven duplication, not on spec.

Naming and composition inside the sheets (#1167)

The previous section settles how many sheets and why; this one settles how a class is named and how a component is composed, so that a proposal framed as "adopt BEM" or "adopt utility-first" is answered from measurement rather than from scratch. Measured over the 28 sheets with the shared parser (tests/Support/CssSelectors.php — the same one the guards read, so the numbers cannot disagree with them), counting a rule per rules() entry, a selector per of() entry and a class per distinct \.([-\w]+) match inside one: 2,121 rules · 2,502 selectors · 1,199 distinct classes, of which 1,094 carry the ffc- prefix and 105 do not. The prefixed share is the figure that moved most since #1167 first took it (1,049 against 147), and it moved because #1170 renamed the unprefixed variants — so re-run the reading rather than quoting it; what does not move is that the unprefixed remainder is almost entirely vendor (CodeMirror, WordPress).

Where each methodology already sits — measured, not aspirational:

Level Evidence
OOCSS partial 108 occurrences (79 distinct) of the compound unit .ffc-a.ffc-b already split object from modifier (.ffc-day.ffc-selected). The skin half is not split: dozens of rules across the sheets declare one background: var(--ffc-X-bg); color: var(--ffc-X-text) pair and nothing else, each under its own name, with more carrying the pair inside a larger component — 57 and 37 when #1168 counted them (#1168)
BEM 5 islands, ~5% 33 __ elements + 24 -- modifiers, all inside ffc-settings-tabs, ffc-form-tabs, ffc-qr-modal, ffc-blocked-roles and the autosave badges
SMACSS state layer yes, layout layer absent 14 is-/has- classes, and every one appears compounded with an ffc- class, never bare — which is why #1152 never saw them, and a bare .is-open would fail it today. Zero l-, zero js-
Atomic / utility-first exists, kept small on purpose ffc-admin-utilities.css. It held 42 classes when this was measured, 9 of them screen components and 3 dead; #1171 moved the components to their owner sheets and deleted the dead ones, #1176 merged the margin utilities a moved step had made duplicates, and UtilityScopeTest now enforces one bare class per rule

Renaming a class is safe only as far as you can find who emits it, and a grep cannot (#1170). tests/Support/CssClassEmitters.php is the scan; CssClassEmissionTest keeps it honest. It knows eleven shapes, and every one of them is in the test because it failed first — a literal attribute, the jQuery/DOM class API, a class used as a selector inside a string ('.ffc-timeslot:not(.ffc-timeslot-full)' — renaming without touching the selector breaks the handler silently), and four ways a name is assembled at runtime: concatenation, an echo embedded in the attribute, a printf placeholder, and PHP interpolation. A name built at runtime exists as a literal NOWHERE, so the scan records the prefix instead and a class counts as emitted when it starts with one. The eighth shape is the one the #1170 renames needed and is easy to miss: a token concatenated onto an attribute that already exists arrives carrying its separator space ('<table class="ffc-appointments-table' + (past ? ' past-appointments' : '')), so a pattern anchored on the quote does not see it — three dashboard classes had no findable emitter until the scan trimmed inside the quotes, and teaching it immediately retired two entries from the open-questions list.

Three more shapes were found by auditing the open-questions list itself, and they are the reason that audit is worth repeating (#1182). Nine of its twenty-four entries had an emitter sitting in the repository — the scan simply could not read the form, which is the category whose written answer is teach the scan, never lower the guard. The first is the mirror image of the separator-space shape above, and shipping one without the other is what hid four dashboard tables for as long as the list existed: where that one loses the token at the END of a concatenation, this one loses the name at the START, because '<table class="ffc-appointments-table' + … never closes its attribute quote and the name arrives with the apostrophe glued to it. The second is several classes in one string — 'ffc-shortcode ffc-form-wrapper ffc-has-geofence' — where the single-token anchoring kept only the runtime prefix of the last name and dropped the whole ones before it; the discriminator that keeps the net closed against wp_enqueue_style handles is requiring two or more tokens, since a handle is always one. The third is a library option whose value is a class: jQuery UI's sortable({ placeholder: 'ffc-sortable-placeholder' }) puts that string on the element it inserts, and the word class appears nowhere near it, so no context window reaches. A fourth site needed only the abbreviation cls added to the context signals.

Two mechanics from that pass. The parity warning below is not theoretical, and the naive fix reproduces it: matching (["'])(.*?)\1 over a window scanned from the file's start had already lost parity by the line that declares ffc-has-geofence, so the multi-class form found nothing until each alternative consumed a whole literal, escapes included. And snapping the context window to line boundaries was measured and removed — it was the intermediate hypothesis for that same failure, and once the literal regex landed it resolved zero classes, which by this file's own rule is churn.

Three things it measured, all of them a trap the first version fell into. Too loose reports nothing wrong: the first run said zero classes were orphaned, because the bare prefix ffc- — from 'ffc-' + Date.now(), which builds an iCal UID — covered all 1,046. A prefix now needs ffc- plus a segment, and the loose forms need the word class nearby, or every wp_enqueue_style handle counts as a class. Matching quoted strings by alternation loses quote parity at the first apostrophe inside a double-quoted string and reads the rest of the file shifted — the same reason CssSelectors has to be quote-aware; the selector form scans the source directly instead. And the output is 26 classes with no known emitter, which is a list of open questions rather than of dead code: each is either genuinely dead, emitted by a shape the scan does not know yet — in which case teach the scan, never add to the list — or applied by something outside our code.

All three answers turned out to have occupants, and the third one is the one nobody expected (#1182). ffc-txt-center and its three siblings were recorded here as the example of genuinely dead — "appear nowhere at all" — and that was true of the repository and false of the world: ffc-pdf-core.css carries a section headed UTILITY CLASSES FOR CERTIFICATE TEMPLATES and two comments reading "Add class ffc-responsive-logo to img tag to enable". They are a published API for the certificate body an administrator writes, which lives in the database, so no code scan can ever find their emitter and deleting them breaks every certificate already using them. The same sheet has a section headed LEGACY CLASSES (Backward compatibility), four of whose classes were never emitted by our code at any point in the repository's history — so what they stay compatible with is also outside it, and they are a shim, subject to "Legacy and tech debt"'s evidence rule rather than to a scan. The list is therefore three lists now: TEMPLATE_API (7, published, not debt), LEGACY_SHIM (4, logged in "Legacy and tech debt"), and WITHOUT_EMITTER (4, still open). The lesson generalises past CSS: read the file before concluding from a grep — the answer was written in the sheet, in a section heading, the whole time.

Three idioms said the same thing; #1170 left two. The measurement found a prefixed modifier (.ffc-day.ffc-selected, 108 occurrences), an unprefixed one (.ffc-consent-status.consent-given, 38 names) and a SMACSS state (.ffc-cap-role.is-on, 12). The rule now is by function, not by word: what the interaction turns on and off is a state and stays unprefixed as is- / has-; everything else of ours — a variant the data dictates, and an element of the component — takes ffc-. tests/Unit/ClassNamingIdiomTest.php refuses the third idiom, with the vendor names that legitimately stay unprefixed listed by family and by name, each with its reason.

It was never a collision risk — all 38 appeared compounded with an ffc- class, which is why #1152's anchor ratchet never saw them. It was a readability one, which by this file's own priority rule is recurring and reader-facing rather than cosmetic. But the pass found a second cost that is not about reading at all: a variant assembled at runtime out of a bare word is invisible to CssClassEmitters by construction, because the scan only records an ffc- prefix — so $progress_color = 'complete' named a class nothing could find. With the prefix the variable holds a literal and the scan finds it, and that is how four dead rules surfaced that the bare name had been hiding: .checkbox-label-text, .ffc-collapsible-section.inline, .ffc-stats-box .stat-value.success/.warning (the cache tab only ever emits info), and .button.loading with the @keyframes only it used. Two more were dead by plain measurement: .ffc-lgpd-consent.error and .checked, whose own comment said "(opcional)".

Two mechanics worth not re-deriving. Rename the token, never the selector: collapsing .ffc-detail-row .value.code into .ffc-detail-value-code drops specificity from (0,4,0) to (0,2,0), which is a different render wearing a rename's clothes — the first attempt did exactly that and was undone. And the proof of zero movement is static and total: every rule on the previous develop, with the rename map applied and the ten dead rules removed, is byte-identical in selector and declaration to what is there now; the second layer is that all 40 new names have a findable emitter, which is what proves both ends of each rename moved together. Of the 147 unprefixed classes only 38 were ever ours — 18 are CodeMirror's, 16 WordPress's, 6 were bare in #1152's baseline — and of those 6, the four that were ours (.delete-link, .status-active, .status-cancelled, .button.loading) left the baseline; .tablenav.top and .tablenav .actions stay, because there what is missing is a page anchor, not a prefix.

The largest finding is not a methodology — it is a scale that nobody reads. --ffc-spacing-* has existed for as long as the palette and has 3 consumers out of 1,238 spacing declarations; 30 distinct px values are in use. This is the typography story before #1148, with one aggravation: the declared ladder (5 · 10 · 15 · 20 · 30, 550 uses) does not contain the ladder the code actually uses (2 · 4 · 6 · 8 · 12 · 16 · 24, 624 uses), and neither dominates. border-radius was half-adopted (82 tokenized against 186 literals, 71 of those literals being the 4px that --ffc-radius-sm already is) and z-index had no scale at all — 1, 2, 1000, 9999, 100000, 100100, 999999, 2147483647, the last being int32's maximum and the classic sign of a stacking war. Both were adopted in #1171, again with zero movement: 149 radius and 16 layer values now read a token. Two things it decided are worth not re-litigating. The four escalation values were deliberately NOT collapsed — 9999, 100001, 999999 and the PDF sheet's 2147483647 are four different answers to "be above the wp-admin bar, which is 99999", and merging them changes the ORDER between elements on the screens where they coexist; that is a behaviour change needing screen-by-screen verification, not a conversion, so they stay literal and visible in the budget. And the nine components sitting in ffc-admin-utilities.css moved out, in the pass that #1171's own comment said they needed. The trap that made it a pass rather than a tidy-up: that sheet is a dependency of ffc-admin-css and is enqueued in exactly ONE place (AdminAssetsManager::enqueue_admin_base_styles(), which enqueues the two in consecutive statements), so everything in it reaches every FFC admin screen — moving a component to its owner sheet narrows its reach, and if that sheet is not enqueued where the component renders, the styles vanish with nothing to catch it. That single enqueue site is also what makes four of the nine reach-identical by construction (ffc-admin.css cannot load without ffc-admin-utilities.css having loaded); the five that genuinely narrowed had their target's gate read against the emitting screen. The proof has to be per screen, with only that screen's sheets — loading all of them hides exactly the defect the move risks: a first run of the harness forgot to substitute the path of the two extra sheets and reported the settings components losing every property, which is precisely what a missing enqueue would look like. With the harness fixed, all three screens render identically, and the 2,498-rule set is byte-identical before and after. UtilityScopeTest now enforces the docblock — one bare class per rule, no descendant, no pseudo-class, no element sub-selector — so the scope is a rule rather than a request.

What this means for the theme arc: a third colour theme costs one line per colour token, because dark is light overridden and the dark block redefines the colour tokens and only the colour tokens — every one except --ffc-text-on-warning, and no non-colour token at all. That is the invariant worth holding, and it is re-measurable: read the two blocks and compare their keys. A density theme (compact/comfortable) — the more likely accessibility ask for public-sector forms — is impossible today, because spacing, radius and z-index are not tokens. That is the same finding seen from the reuse angle, and it is the reason #1169 is ranked first.

Standing decisions, so they are not re-litigated:

  • No BEM across the base. The ffc- prefix already carries the isolation BEM's naming would, over ~1,000 classes; converting is churn against a property that already holds. The 5 BEM islands stay as they are — internally coherent, and rewriting them to the house idiom is churn in the other direction. What survives from the idea is component-specific naming, which #1151 / #1154 / #1162 have been doing one component at a time.
  • No ITCSS with per-area bundles. Measured in the previous section: 1.9×/2.2× per screen, and conditional loading is a safety property rather than an optimisation. Two of the seven layers (generic/reset, elements/bare tags) presuppose owning the document, which a plugin does not.
  • No Tailwind-style utility-first. It needs a CSS build the project does not have (cleancss minifies in place, it does not compile) and a class scan that cannot see what PHP assembles at runtime. The utility layer that does make sense here already exists; the work is keeping it small (#1171), not growing it.
  • No layout layer (l-). Zero occurrences, and the split it would express is already made per sheet.
  • A scale is worth building only where a second consumer of the same value exists. Colour, typography and — since #1169 — spacing have earned theirs. A one-off number, an optical nudge or a hairline stays literal with the reason inline, exactly as the typography section describes.
  • Spacing reads the scale (#1169), and adoption moved nothing. --ffc-spacing-* had 3 consumers out of 1,238 spacing values across 30 distinct px values — values, not declarations, since padding: 12px 16px is one declaration and two values, and counting the other way gives a different number for the same sheets; 1,147 of them read the scale after the pass, and SpacingTokensTest freezes the remainder per sheet, which is the figure a reader should trust because a test asserts it. Four things it measured are worth not re-deriving. A pure 8-point ladder would have moved 14% of declarations by 3px or more — the obvious-looking choice was the destructive one, and the snap cost is what ruled it out; the progression 2 · 4 · 6 · 8 · 10 · 12 · 16 · 20 · 24 keeps 74% exact. Zero movement is provable, not assertable: the 1,238 declarations were resolved back to px before and after and the two lists are byte-identical, and five components were re-rendered against origin/develop in Chromium with no divergence. Two steps were named by their VALUE on purpose, and that is what got them deleted — --ffc-spacing-5 and --ffc-spacing-15 were not steps but what remained of a second ladder (182 declarations), and the naming rule existed so they would leave by deletion rather than by a rename that hides the debt. #1176 removed them: there is one ladder now, 92 literals left of 162, and the movement was bounded and listed — 178 declarations moved 1px, 63 moved 2px, none moved more. Three of the four values were ties on distance, so the tie-break is the typography rule again — relationship, not rounding — read off what the code already does for the same job: gap went UP to xs (its own job already used 6px 21× against 8×), every other property went DOWN to 2xs (4px wins 32× against 17× outside gap), and 15 went to xl on plain distance. 14px and 18px were the exact midpoints of the two widest gaps (12→16 and 16→20) — the issue's own secondary trigger — and they snapped down rather than becoming rungs, because an eleven-rung ladder is a dictionary; the pair padding: 14px 18px became 12px 16px, a shape the sheets already used 9 times. Unlike phase A there is no zero movement here, so the proof is different: every declaration's delta is measured and listed, and Chromium reports zero new horizontal overflow with box shifts of at most 4px. One consequence worth knowing before the next scale change: the margin utilities were named by pixel (.ffc-mt-15, .ffc-ml-18), so moving a step made nine class names lie — they now read the step (.ffc-mt-xl, .ffc-ml-xl), and .ffc-mb-4/.ffc-mb-5 merged because they had become the same value. A utility named by its value re-breaks every time the scale moves. And a sheet enqueued without ffc-common must not read the scale — ffc-pdf-core.css and ffc-code-editor-dark.css are deliberately independent (print is light, the editor theme is dark in both), so a token there would invalidate the whole declaration; the conversion hit both and direction B caught it.

Page scope on the admin wrap (#1184)

Every admin screen the plugin draws says so in its markup. Two classes on the div.wrap, and both have real consumers — this is not one level held in reserve:

  • ffc-admin-page — every admin screen of ours. The anchor for a rule that holds on all of them: ffc-admin.css is enqueued twice under two handles — ffc-admin-css on is_ffc_page() screens and ffc-admin on users.php — so its .tablenav reaches a core screen today. (The CssNamespaceAnchorTest baseline used to say AdminUserColumns enqueues it unconditionally; measured, that method returns unless the hook is users.php.)
  • ffc-page-<slug> — one screen, where <slug> is the screen's ?page= minus the leading ffc-. The anchor for a rule that belongs to one screen and reaches the others through a wide gate: ffc-admin-submissions.css is gated on is_ffc_page(), which matches any ?page=ffc-*, and its .button[title] builds a tooltip on every button with a title on every FFC screen.

A sheet that serves a whole menu family anchors on the family, not on the generic. ffc-audience-admin.css is gated on strpos( $hook, 'ffc-scheduling' ), so its .column-* rules descend from [class*="ffc-page-scheduling-"] — one selector that mirrors the gate, instead of the eleven screen classes it would take to spell out, and at the same specificity as a class. That level is not cosmetic: it is what genuinely resolved the column-actions / column-status pair in StylesheetOwnershipTest::ALLOWED, whose exception read "distinct admin screens" — a reason that lived entirely in the enqueue gate, which that scan cannot read. Anchoring both sheets on the generic ffc-admin-page instead would have made the guard stop reporting the pair without making the collision any less possible: the same risk, now invisible. Anchor at the level the rule's scope actually is.

A screen WordPress draws takes a third anchor, and it costs nothing: the post editor for ffc_self_scheduling has no wrap of ours, so ffc-calendar-editor.css anchors on body.post-type-ffc_self_scheduling — a class core already prints. The same shape as .post-type-ffc_form, which the guard's anchored-forms list already carried.

tests/Unit/AdminPageScopeTest.php enforces it and freezes the screen map. It is deliberately the cheap half — it proves the anchor is in the markup, never that a rule reads it. What charges the reading is CssNamespaceAnchorTest, and that is where the debt shrinks. Splitting the two is what makes the delivery provable by construction: adding a class no rule reads cannot move a pixel, which is the strongest proof of zero movement there is.

The debt it unlocked, and what shrinking it measured (#1184 segunda metade). The CssNamespaceAnchorTest baseline went from 41 entries / 51 occurrences to 3 / 3 (and to zero in #1202): 48 occurrences gained an ancestor, and the rewrite is provably nothing else — stripping the anchors back out reproduces origin/develop byte for byte in all six sheets. Rendered across the 7 screens × 2 themes with only that screen's sheets plus WordPress core's own, 26,082 (element × property) pairs compare identical; the harness's own mutation (one extra declaration on an anchored rule) reports 24, so a zero there is a measurement and not a collapse.

Three things that only the render settled:

  • Narrowing a rule to the screen its docblock names is not always safe, and the split is by what the rule MEANS. ffc-admin-submissions.css loads on every ?page=ffc-*. Its four emoji rules (.button[href*="action=edit"]::before and siblings) are row actions of one screen, so they took ffc-page-submissions — and a scan for the markup they target found it in exactly one file, so that narrowing moves nothing in practice. Its tooltip family (.button[title], :hover::after/::before, :focus) is a generic admin affordance reaching five screens today, so it took ffc-admin-page: narrowing that one to the submissions screen would silently drop tooltips from four others. The rule in the wrong sheet is a #1162 question, not a namespace one.

  • The first harness manufactured a divergence, the same way #1185's did. Synthetic markup put submissions row-action buttons on the audiences screen, and the diff dutifully reported the emoji disappearing — for markup no screen renders. Reading the emitter found class="button" with action=edit|trash|restore in one file only. Build the case from the emitter, not from what the selector suggests.

  • .tablenav { clear: both } restated what WordPress core already declares (wp-list-tables.css l.678 on the 6.4 floor, l.675 on 6.9), which is why narrowing it off users.php measured zero. Our copy won on specificity and declared the identical value, so it was dead on every screen. Deleted in #1202 item 4, together with its sibling .tablenav .actions { overflow: visible } — a different kind of dead: core declares no overflow there at all, and visible is the property's initial value, so the rule only did something if some sheet hid it, and none does.

    The control mutation said more than the removal did, and that is the shape to reuse for a deletion. Removing both moved nothing across 22,908 (element × property) pairs; flipping clear: both to clear: none in place moved exactly two — the two .tablenav elements. So the harness sees that property on those elements, and our rule is genuinely the cascade winner: it is a no-op because the value is identical, not because it loses. A removal that measures zero cannot tell those two apart on its own, and only the second is a safe deletion. The render is screen-independent here because a static scan first established that no sheet, ours or core's, declares clear on .tablenav or overflow on .tablenav .actions anywhere else.

The three that survived that pass were the #tab-* ids of the user dashboard, and they were left alone deliberately, because there the fix is renaming the id, not anchoring it. An id is unique in the document, so a theme with #tab-profile does not repaint — it breaks getElementById, the tabs' aria-controls and the event delegation. Anchoring inside a container would leave the generic id in the document, which is the whole risk. #1202 item 1 did that rename, which is what took the baseline to zero — so nothing is left here; the lesson is, and it is the one to reuse: an anchor is the wrong fix for a name that must be unique.

Three things the measurement found, none of them guessable:

  1. ?page= is not available as a derived source. Two of the fourteen slugs — ffc-scheduling-dashboard and ffc-scheduling-environments — exist as a literal nowhere in the repository: they are built by concatenation (self::MENU_SLUG . '-dashboard'). A check that required finding the slug in the source would report both correct classes as typos. Same finding CssClassEmitters records for a class assembled at runtime, and the reason the map is a frozen register rather than a derivation.
  2. .ffc-settings-wrap was never a competing page-scope name. The issue counted three precedents with no convention; measured, they are three different things. ffc-settings-wrap is on two screens (ffc-settings and ffc-scheduling-settings) — it is a layout class, "this screen uses the vertical-tab design", and it stays. ffc-recruitment-admin is a genuine page anchor with exactly one rule reading it (.wrap.ffc-recruitment-admin .card), and it is the precedent this convention generalises. Only ffc-settings-page was a third name for the page scope, and it is the nested one below.
  3. Two div.wrap were nested inside the Settings wrap — fixed in #1202 item 3, and NESTED is now empty. What the fix measured is worth keeping, because two halves of it look alike and are not. Deleting the six duplicated .ffc-settings-page <descendant> selectors moves nothing (0 divergences across 14,940 pairs): each had a .ffc-settings-wrap sibling in the same rule and the inner div is a descendant of the outer, so the rule already applied. Dropping wrap from the inner div moves the render, and the delta decomposes to a single cause — six margin properties going to zero on that div, and the 22px of width it gains propagating to descendants; no colour, font or display property moves. The result is the point: the user-access tab's first card lands at exactly the x and width of a sibling tab's, where before it was 2px indented and 22px narrower.

Core's .wrap margin doubles horizontally and NOT vertically, which is the opposite of what "nesting doubles the margin" suggests: the inner margin-top: 10px collapses with the panel's, so the card's y is identical in all three states. And the two nested wraps were not the same kind of thing — one was live on every visit to that tab, the other only the fallback for a missing view file that always ships, whose normal path already opened ffc-settings-wrap.

Theme (one palette, one root class — #1126)

Every colour the plugin paints comes from a var(--ffc-*) token declared in assets/css/ffc-common.css. The dark theme is that same sheet's :root.ffc-dark-mode block redefining part of the tokens — the dark theme is the light theme overridden, never a second set, which is why DarkModeCssTest::palette('dark') merges the two before measuring. AssetHelper::enqueue_dark_mode() puts the class on <html>, reading ffc_settings['dark_mode'] (off / on / auto).

The plugin's own setting is the single source of truth. There is ZERO prefers-color-scheme in the CSS — the OS is consulted inside ffc-dark-mode.js, and only when the setting is auto. A rule keyed on the media query would silently outrank the administrator's choice; don't add one.

Seven things the #1126 arc measured that are not guessable, each of which shipped as a defect first:

  1. The palette alone is not the theme. A fully tokenized sheet renders light when the toggle script never reaches the page — every var(--ffc-*) resolves through the light block. Four public screens were in exactly that state.
  2. An undeclared custom property invalidates the WHOLE declaration. It does not fall back to the literal that was there before, so the element renders with no colour at all. Tokenizing a sheet whose enqueue does not depend on ffc-common makes the screen worse. A local :root { --ffc-… } inside a component is the same defect wearing a different hat: it shadows the palette for everything inside, so :root.ffc-dark-mode can never reach it.
  3. Text that declares no color inherits from OUTSIDE this repository — core's body { color: #3c434a } in wp-admin, the active theme on a public page. Both are near-black: 1,28:1 on a dark ground. The palette therefore carries a base pair per root we own (body.wp-admin — never bare body, since the class also lands on public pages).
  4. A form control does not inherit color at all. <button>, <input>, <select> and <textarea> take buttontext / fieldtext from the user agent, so the base pair cannot reach them. A rule that paints a control's ground must paint its text.
  5. opacity on a row fades text and ground together, multiplying down whatever contrast the tokens guaranteed — and the pair meter cannot see it by construction, because the tokens stay correct. Say state with colour, not with opacity. OpacityContrastTest freezes every opacity below 1 in the sheets, per sheet and per selector, each with its reason, a ratchet both ways. It cannot read contrast and never will: what fading does depends on what is BEHIND the element, which is DOM and not stylesheet — what it guarantees is that none enters without somebody having looked. Three categories are legitimate: an inactive component (SC 1.4.3 exempts it, the same reason DERIVED_EXCEPTIONS already uses), a transient state (loading, :hover), and opacity: 0, where the element is not painted at all. The measurement that produced it also found a different defect wearing the same clothes: faded elements that compute color: rgb(0,0,0) because nothing in our sheets declares their colour — they inherit from outside the repository (#1126 defect 3).

That register said four, the list said three, and #1185 re-measured by render: five elements, and not one of them is a defect. All five sit under a root the base pair already names, so in the dark theme their colour is ours; in the light theme they inherit, and that is the design — the base pair is written for the dark theme because that is where inheriting near-black lands at 1,28:1, while on a light ground inheriting is what being a guest in someone else's document means. With the fade included they measure between 5,55:1 and 21:1 in both themes, so the fade stays and the entries now carry the number instead of a promise to look later.

The lesson is about the harness, not the CSS, and it cost two wrong conclusions before it was caught. Measuring those elements in a bare page — no body.wp-admin, no .ffc-shortcode wrapper — reproduces defect 3 perfectly, because the base pair is keyed on exactly those roots. The dashboard calendar "measured" 1,64:1 that way and a fix for it was written and then reverted: with the real DOM (it is an add_submenu_page screen, so body.wp-admin is its root) the number is 10,46:1 and nothing was ever wrong. Read the emitter for the root before believing a contrast measurement — the same rule "Stylesheets and theme" already states for CSS class names, and the same trap as the setContent one, one level up: there the sheets never loaded, here the DOM was missing the one ancestor that does the work.

A static guard cannot replace that render, and a count says why: well over a hundred rules across the sheets paint a surface token without declaring text (#1126 read 122, of which 75 on a single-class selector; a re-run counting every rule with a var(--ffc-*) background and no color gives 240, of which 122 on a bare class — the two readings differ because "a rule that paints a surface" was never pinned down, and that is the point), and nearly all of them are correct because their text descends from a root that IS on the list. What separates a real defect from those is the computed colour under the real ancestors — DOM, not stylesheet. 6. An inline style="" beats the tokenized class. Tokenizing is dead code while an inline colour exists. Where the background is a colour the operator picks, no token can work — the foreground comes from Core\ContrastColor::on(), which computes it by WCAG luminance. Its two candidates are pure black and white on purpose: the palette's near-black #1d2327 drops the worst mid-tone of the RGB cube to 3,99:1, while black holds 4,58:1 over any colour. 7. Categorical hues are not theme colours. The twelve capability-group hues must stay distinguishable from each other, so they stay literal with the reason inline. Same for #adminmenu (it follows the user's own wp-admin colour scheme), vendor brand colours, an editor theme that is dark in both themes, and print.

Five guards, and none was written on spec — each came from a defect that had already shipped:

Guard What it blocks
AdminStylesheetTokensTest A a colour literal per sheet — a ratchet, 0 for the converted ones; counts hex, rgb()/hsl() and the named colour (#1168)
…B / B2 a sheet reading tokens without declaring ffc-common; a method enqueuing the palette without the toggle
…C a var(--ffc-*) nobody declares
…D the base pair exists, paints through a token, and names only live roots; the notice's text nodes likewise
…E a form control given a ground but no text colour
DarkModeCssTest every painted pair against its WCAG floor, both themes (4.5:1 text, 3:1 signal) — two halves, see below

Each carries a self-check that fails when its own scan collapses — an empty result must never read as "clean" (the #1071 / #1094 lesson).

The contrast meter is two halves, and the line above was aspirational until #1168. DarkModeCssTest::pairs() is a hand-written list of token pairs — it measures what somebody remembered to list, and --ffc-danger over --ffc-danger-bg was never on it, shipping at 4,25:1 in the LIGHT theme on a public-CSV warning. So a second half now derives the pairs: every rule that declares color and a background in the same rule is measured in both themes against the 4,5:1 text floor, blocking at zero. Keep both — they cover different things: the scan only sees what one rule declares together, while the pair that inheritance creates (the base text over --ffc-gray-100, a label whose colour comes from its container) is invisible to any static scan and is exactly what the curated list is for.

Three things the derived scan measured, none of them guessable:

  1. A sheet certified "0 literals" was painting white on four dark grounds. The literal ratchet's regex was /#[0-9a-fA-F]{3,8}\b|\brgba?\(|\bhsla?\(/ — a named colour is a word, so it matched nothing, and ffc-admin-submissions.css carried color: white on the PDF, delete and restore buttons plus a tooltip, measuring 1,23 · 2,52 · 2,68 · 2,78:1 in dark mode. The 2,52 is the same number this file already records as the historical on-primary bug: the palette fixed the token, these four rules never adopted it. Direction A now counts named colours, only inside a colour-valued property (font-family: 'Whitney' is not a literal).
  2. Unresolvable must FAIL, never skip. The scan's first version could not read !important or white and reported those four as "unresolvable" — i.e. it would have passed over the worst defects in the repository. A colour the scan cannot resolve now fails the test and asks to be taught, which is the same rule the dbDelta gate states as "never count as clean what it did not look at".
  3. Only two pairs are legitimately below the floor, and both are inactive controls carrying cursor: not-allowed — a full time slot and a readonly/disabled field. SC 1.4.3 exempts text that is part of an inactive component; they sit in DERIVED_EXCEPTIONS with that reason, and a test fails when an exception stops matching any real pair.

The skin object was measured and NOT built (#1168). Consolidating the 57 rules that declare only a --ffc-X-bg / --ffc-X-text pair into six shared classes is the textbook OOCSS split, and it was rejected here: the tokens already deliver "change the colour in one place", so the duplication is syntactic, and migrating the status-badge families means a status→skin map in PHP and JS across ~36 files, because the class is built by concatenation ('ffc-dashboard-status-' . $status). The guard delivers the property that mattered — a wrong pair becomes impossible to ship — without the churn.

Standing decisions, so they are not re-litigated:

  • Print and PDF are light by definition. ffc-pdf-core.css and the certificate preview's white paper are documents that may be printed; they do not follow the theme.
  • Core screens stay light — profile.php, users.php, the post-editor metabox. The content area of every core colour scheme is white, so our fields rendering light there is consistent, not a bug. Listed in TOGGLE_NOT_NEEDED with that reason.
  • The code editor has its own setting, code_editor_theme on the General tab, immediately below Dark Mode: dark (default), light, or auto (follows dark_mode). It is not hard-coded. It sat on Advanced until #1148 and nobody found it there — auto resolves by reading dark_mode, so the two belong side by side. It was moved, not mirrored: SettingsAutosaveFieldPlacementTest refuses one autosave key on two tabs, because each tab's Save rebuilds its own fields and two copies drift.
  • Typography reads the scale (#1148 item 4). Seven steps in ffc-common.css — 2xs 11px · xs 12px · sm 13px · base 14px · lg 16px · xl 18px · 2xl 24px — frozen by TypographyTokensTest, a per-sheet ratchet in the mould of the colour one. Three things it measured that are the opposite of what they look like. The xs step moved from 11px to 12px, and that rename cost nothing because the scale had zero consumers: 415 font-size declarations and not one read a token. 12px had 69 uses — the third most frequent value in the codebase — and the scale skipped it, which is why the sheets invented it. Seventeen of the eighteen "same selector, two sizes" cases are @media, a deliberate step down on phones, not drift; the opposite hypothesis was the intuitive one and was wrong. An icon sized by font-size is not typography — it is a glyph box (dashicons, a &times;, a stat card's icon), and those stay literal with the reason inline, like the categorical hues. The adoption finished at 23 literals in the whole codebase, and the four families that keep one are worth knowing before adding a fifth: a glyph box; a hero number a card sets deliberately above the scale; an em that is relative to its parent because the component is dropped into contexts of different sizes; and a value already one step below the floor inside a phone @media, where rounding up to 2xs would erase the distinction the rule exists to make. 15px is equidistant between base and lg, and the tie is broken by relationship, not by rounding: a heading that lands on body size loses its hierarchy, so headings go up and body text goes down — 17 declarations turned on that one rule.
  • Each step is max(<px floor>, <rem>), and the form is the point (#1157). rem resolves against the DOCUMENT root, which on the frontend belongs to the site's theme: under the common html { font-size: 62.5% } idiom the whole scale shrinks and sm renders at 8.1px — measured in Chromium, not estimated. The px floor removes that; the rem half keeps the browser font-size preference these public-sector forms want. It is not a compromise — it beats rem in the case that breaks and px in the case that matters. The one honest loss is the combined case: under a root-shrinking theme the user's preference is absorbed by the floor until it exceeds it (at 62.5%, asking for 24px still yields the floor's 13px) — still better than the 12.2px plain rem would give there. The two halves must agree at a 16px root — max(13px, 0.8125rem) is an identity, not a range — and TypographyTokensTest checks exactly that, because a mistyped max(13px, 0.75rem) is plausible on both sides and only shows up rendered, in a theme nobody here runs. This replaced the earlier "accepted exposure": the survey of real themes #1157 proposed became unnecessary once the answer stopped depending on which theme is installed.
  • A theme change repaints live. The autosave widget announces ffc:setting-saved on document; ffc-dark-mode.js listens and re-applies, including dropping the OS listener when leaving auto. The widget must not learn which keys repaint — it states the fact, an interested script acts.

Open items and the measurements behind them are in #1148 — among them the typography scale, which already exists and is already semantic, and the trade that makes multiple themes cost more than it looks.


5. Domain conventions

Date / time storage convention

Two categories. Pick the right one when adding a new column or touching an existing one — see #249 for the migration roadmap that retires the mixed pre-#244 patterns.

Category A — Instants (a moment in physical time)

Things like "the user submitted this form at X", "the admin called this candidate at Y", "this audit row was written at Z". Use these rules:

  • Schema: BIGINT UNSIGNED.
  • Write: store time() (PHP) — returns UTC unix seconds by construction, independent of date.timezone or the WP TZ setting. Never current_time('mysql'), which respects WP TZ and produces a string that drifts when the admin changes their site timezone.
  • Read: pass straight to DateFormatter::format_datetime($ts) or wp_date($fmt, $ts). Both apply wp_timezone() to render. Changing the WP TZ re-renders correctly with no data migration.
  • Compare: ints compare directly — WHERE ts > UNIX_TIMESTAMP(NOW()), BETWEEN ? AND ?, ORDER BY ts DESC. No STR_TO_DATE, no FROM_UNIXTIME in the predicate.
  • PHPDoc: @var int Unix UTC timestamp (seconds since epoch).

Existing examples: the Public Operator Access audit ring buffer (entry['ts']) was unix int from day one — that's why the only TZ bug that ever appeared there (#247) was a rendering choice (gmdate → wp_date), never a storage issue.

Category A exception — housekeeping timestamps

created_at / updated_at columns that are (1) MySQL auto-managed via DEFAULT CURRENT_TIMESTAMP (and ON UPDATE CURRENT_TIMESTAMP), or (2) PHP-managed but never rendered to end users, stay as DATETIME.

Rationale: these are audit / sort columns only — they never reach a display path that would surface TZ drift. BIGINT UNSIGNED would force PHP responsibility for every INSERT/UPDATE site (MySQL cannot DEFAULT CURRENT_TIMESTAMP on BIGINT) with no user-facing benefit.

Current inventory (point-in-time snapshot — re-verify a column against the live schema before relying on its row here):

Table Pattern Notes
ffc_reregistration_submissions P1 (MySQL auto) ORDER BY created_at in repository
ffc_recruitment_* (6 tables) P2 (PHP-managed, NOT NULL) written via current_time('mysql')
ffc_audience_* (5 tables) P1 (MySQL auto) —
ffc_rate_limit_* (3 tables) P1 (MySQL auto) —
ffc_custom_fields* (3 tables) P1 (MySQL auto) —
ffc_short_urls P2 (PHP-managed, NOT NULL) table name is ffc_short_urls (not ffc_url_shortener, which is only an option-key/meta-box id prefix)
ffc_self_scheduling_* P3 (hybrid: created_at auto, updated_at PHP) —
ffc_activity_log P2 (PHP-managed, NOT NULL) —

If a future feature renders one of these columns to a user, that column must be migrated to Category A storage at that point — not left as a hidden TZ-drift trap.

Category B — Wall-clock (a human commitment, no TZ semantics)

Things like "the appointment is on May 20", "the doctor sees the patient at 09:00". The value means the same thing if the user travels or the server changes TZ — converting to UTC introduces DST/ambiguity bugs.

  • Schema: DATE, TIME, or DATETIME (the combined form). Stored literally, no conversion at read or write.
  • Write: store what the user picked — '2026-05-20', '09:00:00'.
  • Read: render via DateFormatter::format_date() / format_time() / format_datetime() directly. wp_date() applies the site TZ when given a unix int, but with a DATE / TIME string source you feed the value as-is.
  • PHPDoc: @var string Wall-clock DATE in 'Y-m-d' (no timezone semantics).

Existing examples: appointment_date (DATE), start_time / end_time (TIME), date_to_assume (DATE), time_to_assume (TIME).

Always

  • Display goes through DateFormatter::format_*(). No gmdate(), no date_i18n(), no wp_date() outside the helper unless there's a documented reason (e.g. building an iCal DTSTAMP per RFC 5545).
  • Filenames / log keys / API contracts that need a stable ISO format may use gmdate('Y-m-d\TH:i:s\Z', $ts) — but the column those filenames represent should still follow Category A or B.

Settings reads

Read ffc_settings via FreeFormCertificate\Settings\SettingsReader, not get_option('ffc_settings') directly:

  • Use typed accessors when one exists (SettingsReader::emails_disabled(), SettingsReader::activity_log_retention_days(), etc.).
  • Fall back to SettingsReader::get($key, $default) for keys without a dedicated typed accessor.
  • Use SettingsReader::all() when a caller reads 5+ keys (SMTP block, DateFormatter format catalog) and array-style access stays clearer than repeated method calls.

The debug-area toggles (the Debug::AREA_* constants own the list) continue to be read via Debug::is_enabled($area) — that helper has the canonical function_exists('get_option') defensive check and is the typed reader for that subset.

Classes that already encapsulate get_option('ffc_settings') in their own private helper (e.g. UrlShortenerService::get_settings()) do NOT need to migrate — they're already centralized.

Settings writes — autosave toggles (the no-clobber invariant)

On/off toggle switches (.ffc-toggle, rendered via AdminUI::render_toggle()) auto-save on flip: mark the input data-ffc-autosave-key="<key>", add the key to SettingsAjaxEndpoint::allowlist() ({option, path?, type, cap}), and have the tab call enqueue_autosave_infra(). This is the standard for atomic, side-effect-free boolean settings. Two autosave endpoints exist and stay separate on purpose (do NOT unify — bad-façade trap): SettingsAjaxEndpoint (ffc_update_setting) writes WP options under a global cap; FormMetaAjaxEndpoint (ffc_update_form_meta, attribute data-ffc-autosave-form-key) writes per-post meta gated on edit_post of that exact form. The JS is one widget serving both (#1116) — FFC.Admin.autoSaveField in assets/js/ffc-admin-autosave.js, wired by bootAutoSaveFields(), which scans both attributes and takes action / payload / request / strings per endpoint. The endpoints stay separate; only their client did not need to be. That statement was written here before it was true: the form-meta half was a second implementation inside assets/js/ffc-admin.js (delegated, change only, no debounce, its own badge whose CSS hardcoded hex and therefore ignored dark mode), and the cost was not theoretical — the #1114 validity guard had to be written twice to reach both halves. So a change to autosave behaviour is now one edit; a screen that renders either attribute must enqueue ffc-admin-autosave (settings tabs get it from enqueue_autosave_infra(), the form editor enqueues it directly), or its fields silently have no listener.

Two things the convergence measured, worth not re-deriving. Every one of the 16 form-meta fields is a checkbox from AdminUI::render_toggle(), so adopting the settings half's change+input binding changed nothing visible — there is no text field on that path to "save while typing", and a checkbox's two events coalesce in the debounce. And the delegated binding was the half not kept, against the original proposal: nothing exercises it (all fields of both kinds are server-rendered at page load), and adopting it would have meant rewriting the half with 115 fields, per-field debounce, multi-checkbox groups and confirmOff to gain that. bootAutoSaveFields() is idempotent and exported, so DOM inserted later has a supported answer — call it again.

Invariant that keeps autosave and the form Save from clobbering each other — hold it whenever you add an autosave toggle:

  • The autosave toggle must ALSO be a named field inside its tab's <form>, so the form's rebuild-save writes the toggle's current (auto-saved) state, not a stale/absent one. An autosave-only key that a rebuild-save omits gets silently reset on the next Save.
  • Tabs that share ffc_settings save through the central SettingsSaveHandler, which is merge-based ($clean = get_option(...)) and gates each section on the hidden _ffc_tab marker — so saving tab X only rebuilds tab X's toggles, preserving every other tab's auto-saved keys. Keep both properties (merge base + _ffc_tab guard) when touching that handler.
  • Tabs with their own option (geolocation, rate-limit, user-access, ip-diagnostics) full-rebuild that option from their own form; every toggle in the option must be read back from $_POST there.

Admin number inputs — required, and the three times it is wrong

Every <input type="number"> in the admin carries required, enforced by tests/Unit/RequiredNumericInputTest.php — an unlisted one without it fails, and a listed exception that gained it (or vanished) fails too. The reason is the #1114 class: a cleared number field posts the empty string, absint( '' ) is 0, and what that 0 means is per-consumer and was never uniform — on the Rate Limit tab it read as "limit already reached" and barred every submission from every address; on a reregistration campaign it collapsed the reminder onto the campaign's last day. The attribute is only the cheap half: the save handler still has to treat empty as "not supplied", because the browser is not the guard (RequestInput::get_post_string( $k, '' ) then an explicit '' === … branch; ArrayValue::int() already does this, since is_numeric( '' ) is false).

Three shapes make required a bug rather than a missing safeguard, and each is in the guard's KNOWN_OPTIONAL with its reason:

  1. Empty is a documented value. The device-limit and public-CSV fields say "Inherit from global" and their save handler deletes the meta so the read side falls back; the audience calendar's says "Leave empty for no limit". required removes the only way to express it.
  2. The control is not a form field. ffc_qty_codes has no name — it is the argument to an AJAX button. A control without a name is still a candidate for constraint validation, so required there blocks the surrounding form while never being submitted. Same for an empty "add a new row" line inside a shared <form> (ffc_location_new[…], caught in #1115 before it shipped).
  3. Requiredness belongs to the field's definition. User-defined custom fields carry is_required; the markup must emit the attribute from that flag, never hardcode it. Both renderers do (#1120 closed the user-profile side, where the flag had only ever drawn an asterisk). Two types stay out on purpose: a checkbox would have to be ticked to satisfy required, a different promise from "fill this in"; and a value posted through a hidden input is barred from constraint validation outright, so marking it required does nothing.

A required field inside a JS-hidden block blocks the submit against a control nobody can see — constraint validation ignores visibility, and only disabled bars a control (which would drop it from the POST). So the attribute travels with the block: FFC.setRequiredWithin( $container, visible ) strips it and puts it back, using a marker attribute so a re-show never promotes fields that were never required. Call it wherever a block is shown or hidden — the calendar editor's four blocks, the quiz rows, and the collapsible audience sections on the user-profile screen all do. This was not hypothetical: the working-hours rows had carried required inside .ffc-regular-only since 4.1.0, so a stored row with an empty time jammed the save of any calendar switched to custom mode.

Capability naming

All FFC capabilities follow one grammar (ratified in #488, applied plugin-wide):

ffc_<action>_[own_]<domain>[_<qualifier>]
  • Actions (closed vocabulary): view (read-only) · manage (read-write: create/edit/delete/configure) · export · import · edit (modify existing records — narrower than manage) · delete. Flow-specific verbs: book, cancel, download, call, bypass.
  • own_ marks a self-scoped end-user cap (frontend; the user's own data).
  • Domains (canonical): certificates, appointments, audiences, reregistration, custom_fields, activity_log, settings, recruitment, url_shortener, forms_api.
  • Qualifiers: _pii, _settings, _reasons, _history, _smtp, _dangerzone. The last two carve the two most sensitive Settings surfaces out of the blanket ffc_manage_settings (#711): ffc_manage_settings_smtp gates the SMTP transport + Email Model save, and ffc_manage_settings_dangerzone gates every destructive maintenance action (delete-all, cleanups, public-access disabler, submission-link audit, migration execution). A dedicated ffc_export_activity_log (export tier) was also split out of the read-only ffc_view_activity_log so a view-only operator cannot bulk-extract the audit trail. All three ship with one-shot grant migrations that seed the new cap onto current holders, so no one loses access on upgrade.

Settings-write auth — WP-standard, no engine (deliberate, #711). Settings-write authorization uses the WordPress-standard inline pattern: native nonce funcs (wp_verify_nonce / check_admin_referer / check_ajax_referer) + the existing capability chokepoint Capabilities::current_user_can_admin_or(). Do not wrap this in a settings persistence/authorize "engine" — one was tried and removed as net-negative indirection (it fit none of the real writers and only hid the nonce from the WPCS sniff).

3-state permission model

Every admin domain exposes a view/manage pair so each surface has three states — não vê / só vê / vê e edita — with the WP admin (manage_options) above all:

canView = current_user_can('manage_options') || view_cap || manage_cap
canEdit = current_user_can('manage_options') || manage_cap

A manage role does not need to also carry the view cap — canView already includes manage. Hidden when neither; read-only render (disabled inputs, no save, row/bulk actions hidden) when only view. Use Capabilities::current_user_can_admin_or($cap) for inline gates; menu/tab caps take the slug directly (admins hold every FFC admin cap via the activation/ensure_admin_capabilities grant).

Registry, catalog, migration

  • Machine list: CapabilityManager (*_CAPABILITIES consts + module_roles_definition()).
  • Human metadata: CapabilityCatalog::groups(). Invariant (enforced by CapabilityCatalogTest): CapabilityCatalog::all_slugs() must equal CapabilityManager::get_all_capabilities() as a set — adding a cap to one without the other fails CI.
  • Renames ship with a one-shot, option-flagged migration that rewrites grants on every user (user_meta) and every role definition (see CapabilityMigrator::migrate_taxonomy_renames() + Loader::ensure_taxonomy_renamed(); the one-shot migrations live in CapabilityMigrator and role lifecycle in RoleRegistrar since #563 Sprint 2). Renames are a breaking change for external integrations referencing old slugs — call it out in the CHANGELOG.

Security & PII conventions

A full security audit confirmed these hold plugin-wide — keep them that way (the #596 IP-hash fix was the only gap found):

  • Output escaping. Escape every echoed value at the output point (esc_html / esc_attr / esc_url / esc_textarea, or wp_kses_post for rich HTML). The WPCS EscapeOutput gate enforces it; a phpcs:ignore must be justified (the value is provably pre-escaped). DB-stored data rendered on a higher-privileged screen is still untrusted — escape it (see the #564 stored-XSS).
  • SQL. All queries go through $wpdb->prepare() with %d / %s / %i (identifiers), esc_like() for LIKE, and an allowlist (or sanitize_sql_orderby()) for any request-derived ORDER BY / column. Never interpolate request data into SQL.
  • Request entry points. Every AJAX / REST / admin_post handler that mutates or exposes PII checks both a capability (current_user_can / Capabilities::current_user_can_* / a non-__return_true permission_callback) and a CSRF nonce (check_ajax_referer / wp_verify_nonce / check_admin_referer). Destructive actions take a narrower cap (e.g. ffc_delete_*). For per-user data, derive the user from get_current_user_id() and gate any viewAsUserId-style override on manage_options — never trust a request-supplied user/owner id (IDOR).
  • PII at rest & display. CPF / RF / email are stored via Encryption (AES-256-CBC, per-record CSPRNG IV, encrypt-then-HMAC, hash_equals verify); searchable copies use a salted hash. Display goes through DocumentFormatter masked unless a PII cap is held. Tokens use random_bytes / wp_generate_password, never rand/mt_rand.
  • Never log raw PII. Debug logs hash IPs / CPF — a truncated hash( 'sha256', … ), never the value — and stay behind the off-by-default debug toggles regardless. See IpGeolocation and ClientIpResolver (16 hex chars) and PreflightTelemetry (12, salted with wp_salt( 'auth' )) (#596). The truncation length is the call site's, not a plugin-wide constant; this file used to quote one length for all of them and disagreed with PreflightTelemetry for as long as it did. What is plugin-wide is the rule: hash before logging.
  • Outbound HTTP. Validate any request-derived IP/URL before wp_remote_* (e.g. FILTER_FLAG_NO_PRIV_RANGE | FILTER_FLAG_NO_RES_RANGE to block SSRF); prefer wp_safe_redirect for redirects.
  • User-deletion integrity is two layers over a deliberate subset, not every user-reference column. The primary is the app-layer UserCleanup hook (deleted_user), which handles user_id on the user-owned tables: SET NULL (retain the row, drop the link) for submissions / appointments / activity_log (the last via ActivityLogQuery::redact_user_id), DELETE (purge the row) for the pure-relationship audience tables (audience_members / audience_booking_users / audience_schedule_permissions) and user_profiles. MigrationForeignKeys adds a DB-level user_id → wp_users FK backstop on those same tables — ON DELETE SET NULL where the hook nulls, ON DELETE CASCADE where it deletes, plus activity_log (FK-only) — so integrity holds even if the hook is bypassed. This FK migration is not legacy: dbDelta() cannot emit FKs, so it is the sole path that creates them (absent on a fresh install until it runs), guarded by ffc_foreign_keys_db_version. Beyond user_id (#834): the hook also runs an authorship sweep (UserCleanup::anonymize_authorship()) that SET-NULLs the nullable actor-attribution columns (submissions.edited_by; appointments.approved_by/cancelled_by; self_scheduling_calendars.created_by/updated_by; blocked_dates.created_by; audience_bookings.cancelled_by; reregistration_submissions.reviewed_by; recruitment_call.cancelled_by; short_urls.created_by) and the promoted-candidate link recruitment_candidate.user_id. The sweep runs on every deletion (an actor/candidate need not have data-subject footprint, so it sits outside the #322 footprint short-circuit that still gates the user_id blocks) — cost is extra unindexed UPDATEs on the cold deletion path. It is app-layer only: no FK backstop is added for attribution columns (they are low-risk actor refs, not primary user-owned records), and the NOT NULL attribution columns are left as accepted orphans (see the gaps note).
    • The governing principle (apply it when adding any user-linked table): retain records, drop relationships. A row that is a record (a submission, an appointment, an audit line, a reregistration) is retained and its user link anonymised — SET NULL when the column is nullable, or left as an accepted orphan when it is NOT NULL and the row is a retained/legal record (the activity_log precedent). A row that is a pure relationship (a membership, a permission grant) is deleted with the user. Both layers key only on user_id — see the gap inventory below before assuming a new column is covered.
    • Recruitment candidate: ffc_recruitment_candidate.user_id is DEFAULT NULL (a candidate isn't a WP user until promotion; documented in its activator). Pre-promotion there's nothing to null; post-promotion the link is now SET NULL by the authorship sweep (#834) — the candidacy record + its PII are retained, only the WP-user link drops.
    • ffc_reregistration_submissions — accepted orphan (#822, decided). Its user_id (NOT NULL) is covered by neither layer on purpose: a reregistration is a retained record, and the identifying PII lives in the data JSON body (partly Encryption-encrypted, partly plaintext), so nulling the FK alone would be a half-measure. The row is retained; the user_id is an accepted orphan; JOIN consumers degrade to "—". There is no separate custom-field value table — those values are the data column (and wp_usermeta under ffc_%, which WP core + the manual PrivacyErasers clean, but the automatic deleted_user hook does not).
    • #834 resolution (was "known remaining gaps"): (1) attribution columns — the nullable ones are now nulled by the authorship sweep above; the six NOT NULL ones (audiences/audience_schedules/audience_holidays/audience_bookings.created_by, reregistrations.created_by, recruitment_call.created_by) stay accepted orphans (a plain SET NULL would be rejected; they're retained-record actor refs) — nulling them would need an ALTER … MODIFY … DEFAULT NULL first, not worth it. (2) recruitment_candidate.user_id post-promotion — now covered (SET NULL). (3) Row-body PII on account deletion — decided policy, not a gap: the automatic deleted_user hook anonymises the link (user_id + nullable attribution) but deliberately does not scrub the encrypted PII columns in the row body (*_encrypted/*_hash). Account deletion ≠ an erasure request: the manual PrivacyErasers (WP GDPR eraser) is the path that additionally clears email_encrypted/cpf_encrypted/rf_encrypted/… so certificates/appointments stay verifiable records until a subject actually exercises erasure. The only residual is the app-layer-only nature of the attribution sweep (no FK backstop) — an accepted design choice, not a follow-up.

6. Legacy and tech debt

Legacy compat shims — audit log

Inventory of the legacy compatibility shims that remain in the code by design (snapshot — re-confirm the location in code before removing; paths and lines change with every refactor, so the table cites files/methods, never line numbers). Removing them requires evidence that no production installation depends on them.

The two shims previously tracked here — the ensure_legacy_caps_renamed() v1 pre-6.2.0 cap-rename migration and the pre-4.6.15 orphan-cron cleanup (activator + deactivator/uninstall) — were removed in 6.18.0 (#809) as scheduled.

Shim · ffc-pdf-core.css sections 7 and 13 — the four legacy PDF-stage classes (ffc-pdf-stage, ffc-pdf-bg-img, ffc-pdf-user-content, ffc-pdf-temp-wrapper). Risk if removed: a certificate body saved before the rename renders unstyled — each has a live sibling the JS emits today (wrapper, bg, content, temp-container), so the pairs name the rename. Why it stays: the sheet declares section 13 "Backward compatibility", and git log -S over assets/js, includes and templates finds no emitter at any point in the repository's history — meaning what depends on them is admin-authored HTML in the database, which is install state and not visible to any scan. Exit condition: a migration or a report that counts stored certificate bodies still carrying these classes, reading 0 pending on a real install — the cpf_rf_encrypted shape, never a code scan. Logged in #1182; frozen in CssClassEmissionTest::LEGACY_SHIM, which is what stops them from being read as dead code again.

The html/ layout fallback (#865 phase-4) was removed in 6.23.0 (#1087), which emptied this inventory until the above was logged. The whole html/ directory went with it, because both of its exit conditions were met and confirmed on a real install: import_legacy_templates reads 0 pending, and so does rewrite_html_image_refs — the migration that side-loads html/*.png into the Media Library and was the only reason the images had to stay.

The chain that makes deleting the images safe is worth keeping, because it is not obvious. That migration deliberately skips shipped defaults (a stale ref on a default is repaired by re-seeding, never by rewriting it to an uploads URL), so it alone would not have covered them. What covers them is CertTemplateSeeder::maybe_seed(): on any version bump it calls restore(), which refreshes each existing default's body to the shipped source — and SEED_VERSION 2 was precisely the #871 fix that moved those bodies off html/ and onto assets/. Defaults by re-seed, everything else by the migration; between the two, nothing points at html/ any more.

Two related things to know if this ever comes up again. The seeder reads templates/certificate-defaults/, not html/ — had it read the latter, deleting those three files would have broken seeding on every fresh install, which is the very condition that authorised the removal. And both migrations degrade without the directory rather than fatal: glob() on a missing path returns nothing (status "100% complete"), and the rewrite records a "Missing html/ file" error per target instead of throwing.

When a new shim is added, log it here (Shim · Location · Risk if removed · Why it stays), and when a new feature makes one unsafe or inadequate, open a specific sub-issue + a breaking-change banner in the CHANGELOG.

Deprecation cycles — @removal is the date, and it fires (#1309)

A shim in the log above is retired by evidence; a deprecation cycle is retired by a release, and that release has to be written somewhere a machine reads. None of the WordPress deprecation functions carries it — apply_filters_deprecated( $hook, $args, '6.26.0', $replacement ) and _deprecated_function( __METHOD__, '6.25.0' ) both take the version the thing was deprecated in, which is the string WordPress prints to whoever is still listening. The removal release has no field at all, so before #1309 it lived in prose and depended on somebody re-reading the issue at the right moment.

Write @removal X.Y.Z on the surviving code — a docblock tag on the method or the documented hook, a // @removal X.Y.Z -- <why>. line where the thing is a bare registration with no docblock of its own. DeprecationDueTest reads it and fails when FFC_VERSION reaches that release, naming every site; a second direction refuses any apply_filters_deprecated() / _deprecated_*() call that carries no marker, which is what stops the first from measuring only what somebody remembered to mark. Deferring a cycle on purpose means moving the date in the file — that is the decision, and the code is where it is worth arguing.

The prose must not restate the number. The marker is the one place; a sentence saying the same thing is the #1261 class, and one such sentence lives in a translated string on the hooks documentation page, which is why that line carries its own marker.

And prose cannot be the source, which cost the first implementation. It read the until X idiom that was already in the tree and reported two Security docblocks as overdue — both past-tense history about the $token argument that was removed in 6.24.0 (#1048). until 6.24.0 is a live promise or a finished story depending on the tense of a verb several clauses away, and the classifier that tells those apart is the "dictionary describes itself" trap CommentLanguageTest records. History in the past tense is correct and stays; it simply is not a field.

Two mechanics worth not re-deriving. The scan reads comment tokens, never raw lines, so a version inside a string literal or a piece of markup can never be mistaken for a marker. And it resolves T_NAME_FULLY_QUALIFIED beside T_STRING, because \_deprecated_function( is one token carrying the backslash — the #1284 trap, pinned here by a canary over synthetic source rather than by what the tree happens to contain.

Resolved exemplar — how the html/ fallback was retired (evidence, not a deprecation cycle)

Kept because the method generalises, and because what it got wrong the first time is the useful part.

Retiring it was evidence-gated (the cpf_rf_encrypted shape below), not a versioned deprecation cycle. The cycle exists for surfaces whose consumers a code scan cannot see — a public method an external integration might call. Both render sites were internal, with no hook, no filter and no external caller, so nothing invisible could depend on them; what they depended on was install state, which is observable.

Two conditions were written down, and they are genuinely different — conflating them is the mistake that nearly retired the shim early:

  1. The pool seeds on every install — the fallback's own written condition, and the one that was not met before 6.22.0. The CertTemplateSeeder::pool_has_defaults() retry is what made it hold.
  2. Settings → Migrations → import_legacy_templates reads 0 pending — every file an admin dropped into html/ has been imported. It had read 0 in production for some time, but it measures imports, not seeding, and on its own said nothing about condition 1.

The seeder fix and the removal shipped in different releases (6.22.0 → 6.23.0), so an install with an empty pool received the repair, and was observed to have taken it, before losing the safety net.

The lesson worth carrying: the written conditions were incomplete. Measuring the surface at removal time found four sites reading html/, not the two the inventory named, and a third exit condition nobody had written down — rewrite_html_image_refs side-loads html/*.png into the Media Library by reading them off disk, so the images could not go until that migration also read 0 pending. The first removal commit therefore deleted the code and the three .html files and kept the images; they went in a second commit, once that condition was attested too. So: re-measure the surface before retiring a shim, and do not trust the inventory's own list of conditions to be exhaustive — it records what was known when the shim was logged, not what is true when it is removed.

Gathering the evidence to remove a High-risk shim

Not every High shim is provable by data. Before proposing to build a diagnostic, check whether the evidence already exists — the resolved cpf_rf_encrypted case below is the exemplar:

  • Resolved — success/fail keys of get_audit_log_summary() (removed in 6.17.0, #730). They were not provable by any diagnostic (a public static method with no hook/filter/DB trace; an external consumer reading ['success']/['fail'] is invisible to any read-only query). Internally they had zero consumers (the metabox migrated to access_success/failed_access), so they were de-risked exactly as prescribed — a versioned deprecation cycle (announced 6.15.0) + a ⚠ breaking-change banner in the CHANGELOG, removed at the 2nd feature release after the notice — never via an evidence scan. count was NOT deprecated (the metabox reads it) and stays.
  • Resolved exemplar — cpf_rf_encrypted (removed after production read 0 pending). Its evidence already existed in the UI: the split_cpf_rf migration card (Settings → Migrations) shows Pending = rows still carrying cpf_rf_hash in ffc_submissions / ffc_self_scheduling_appointments (CpfRfSplitMigrationStrategy::count_table_status(); a dropped column ⇒ 100% complete). Once a production install read 0 pending, the legacy cpf_rf_encrypted reads (the PDF appointment fallback + its two REST-controller siblings) had no live dependent and were removed. The equivalence that backed the reading: Encryption always writes hash + ciphertext together and the migration nulls cpf_rf/cpf_rf_encrypted/cpf_rf_hash atomically per row, so "pending by cpf_rf_hash" ⟺ "pending by cpf_rf_encrypted". The lesson: a parallel counter keyed on cpf_rf_encrypted would have been redundant indirection — the migration card was already the signal (the façade trap that narrows nothing). (The combined cpf_rf view field is separate legacy and its retirement is tracked apart from the shim.)

Parked architecture decisions

  • Import consolidation (recruitment ↔ audience) — parked / not planned (#788). The #772 export contract deliberately does not extend to import: import↔export share only cheap scaffolding (uuid + {processed,total,done} + TTL), while store (temp file vs staging DB), arity (3 vs 4 phases) and direction all diverge — unifying would be the bad-façade trap. One batched importer (recruitment CsvStagingService) exists today; audience is single-request. Revisit only on a trigger — real audience-import timeouts in production, or a 3rd server-side import duplicating the stage-all→validate→promote shape — then extract from the proven duplication (the #772 way, which factored BatchedCsvExport out of two existing clones), not on spec. Full audit + verdict + lessons in #788.
  • Client-IP default flip (legacy → secure) — parked / not planned (#902, epic #899). The consolidated Core\ClientIpResolver ships the trusted-proxy + Cloudflare-autodetect decision tree, but only as an admin opt-in (the secure strategy on the IP Diagnostics tab), with a notice strongly recommending secure + Cloudflare while legacy stays effective (#908). Forcing the effective default to secure was deliberately not done: the non-IP defence layers (device fingerprint, nonce + capability + rate-limit per user_id) already carry the load, so a stronger IP signal is a quality improvement, not the only wall — and a forced flip risks collapsing rate-limit/geofence into one bucket on shared hosts (the plugin's typical deployment). The shadow-divergence gate that would have justified the flip (ffc_ip_shadow_logging) is opt-in and off by default, so it never accumulates data on its own; keeping #902 open behind it would defer forever by inertia. Revisit only on a trigger — concrete evidence of an exploitable IP spoof (rate-limit/geofence actually bypassed via a forged header), or a consumer that genuinely needs the strong IP signal — then extract the flip from the proven need, not on spec (the #788 criterion). Full rationale in #902 / #899.
  • Declarative settings registry — parked / not planned (#993). A single ffc_settings key is known in up to six places (Settings::get_default_settings(), the SettingsSaveHandler sanitiser, SettingsAjaxEndpoint::allowlist(), a SettingsReader typed accessor, every read-site default literal, the tab view), and nothing forces them to agree. The proposal is one declaration per key — { default, type, sanitize, cap, autosave } — that every consumer derives from, making divergence unrepresentable rather than merely detectable. Deliberately not built: it touches Settings, SettingsReader, SettingsSaveHandler, SettingsAjaxEndpoint, ~17 read sites and probably the module-boundary baseline (days, not hours), which is the bad-façade trap unless the duplication is actually biting. The cheaper guard landed instead and is what the 6.20.1 CHANGELOG entry citing #993 refers to: the 5 measured divergences across 3 keys were aligned and tests/Unit/SettingsDefaultsTest.php now fails CI when a read-site default disagrees with the declared one — so do not read that entry as the issue being resolved. Revisit only on a trigger — a new divergence the guard catches that cannot be fixed by aligning a literal (two consumers legitimately needing different defaults for one key), a third consumer of the key list (uninstall + the autosave allowlist are the two today), or one of the 17 keys read through SettingsReader but never declared producing a #936-style dormant feature. Full design, measurements and verdict in #993.
  • Vendored runtime bundles vs. Dependabot — parked / not planned (#1208). The four hand-vendored bundles in libs/js/ — ALTCHA, thumbmark, html2canvas, jsPDF — are enqueued at runtime and inside the distributable zip, and none is in any manifest, so a CVE in one will never open a PR here. The contrast is the sharp part and it is mechanical rather than a judgement: composer.json's require is php alone and package.json's dependencies is empty, so everything Dependabot can see is dev/CI and never ships, and everything that ships is invisible to it — which is also why the Dependabot section's runtime-hotfix exception is unreachable through Dependabot itself. Three ways out, none obviously right: audit by hand (today; #1203's filename-carries-the-version stamp is what makes it cheap), declare the four in package.json purely so Dependabot sees them — the alert would arrive but the PR would touch only package-lock.json, so a green PR reads as "already updated" when nothing was, — or build them from npm, which actually solves it and adds the build step the project deliberately does not have (cleancss/terser minify in place, they do not compile) while breaking byte-for-byte verification against upstream, the reason to vendor at all. That the manual audit is real work and not a placeholder was measured: the 2026-09-20 pass over the four is what found #1357. One decision is recorded and stands: thumbmark stays at 1.10.1. Measured against upstream source, not the minified build — 1.11.0 adds src/utils/experimental.ts, which fetches a remote payload and runs it through indirect eval, and the gate is not the one its name suggests: options.experimental defaults to false and is read nowhere in src/, while getExperimentalPayload()'s caller logThumbmarkData() is unconditional, so only logging: false closes the path. The plugin already sets that, with a test pinning the call — but a line protected by a test is more fragile than an absent path, and not upgrading costs nothing. Revisit only on a trigger, not on "a new version came out": a published CVE in any of the four, or a fifth bundle landing in libs/js/ — at which point the manual audit starts compounding and option 2 or 3 gets decided rather than deferred. The cheaper guard landed instead and is what closed the issue: tests/Unit/VendoredBundleInventoryTest.php freezes the directory's contents, so the fifth-bundle trigger fails CI instead of depending on somebody noticing, and holds the #1203 shape (filename ↔ constant, no version literal in a libs/js/ enqueue — the defect that had the appointment receipt announcing ?ver=2.5.1 for a 4.2.1 bundle). It cannot see a CVE or verify bytes against upstream; what it guarantees is that the inventory cannot change without somebody saying so. Full audit + verdict in #1208.