Every rule in the README came from a maintainer's verdict on a real pull request. This is the ledger. Quotes are verbatim; the rule column is what the verdict became. Newest first.
| Date | Pull request | What the maintainer said | What it became |
|---|---|---|---|
| 2026-10-09 | PaloAltoNetworks/docusaurus-openapi-docs #1624, #1625, #1626 | Jeff, after three merges shipped with the fork branches still up: "you didn't clean up the branches? wtf"; "I don't want normal known activity processes not fully in our documented model" | sweep records the fork branch of every merged or closed pull request and prints a TIDY line each run until hygiene has deleted it; a routine step lives in the tool's output, not in the operator's memory |
| 2026-09-25 | hashicorp/hcl #840 and #841 | The issue offered a fix without linking one: "I have a fix ... say the word and I will send it as a PR." The next day the author of the series it touched answered "as author of #810 and #811, I welcome a PR for this," and the original hclwrite author added the case "ought to work." The pull request went out a day later than it should have. | The libsrtp rule of 2026-09-09 had no enforcement on the issue route, because issues never pass through file-pr.sh. Issue #26: an issue-filing gate that refuses a body offering a fix, a patch or a PR without a branch or commit URL on the author's fork. Until it ships, no issue leaves the machine with a fix offered and no link. |
| 2026-09-23 | canonical/pebble #938 | Closed with "Thank-you for the PR. The fix has been absorbed into #1075." A second maintainer had cherry-picked the commit into their own pull request, which merged to master with a Co-authored-by trailer naming the author. | A closed pull request is not a loss until the closing comment and the closing commit have been read. When a maintainer absorbs the fix, count it as merged by that route, link both pull requests in the receipt, and delete the fork branch as for any merge. |
| 2026-09-21 | square/kotlinpoet #2354 and #2391 | Closing another contributor's pull request: "There was supposed to be a box in the PR description that you should've ticked, but you deleted it." and "Closing as the CLA was not signed and there was no response." The same day the contributing guide gained: "please do not delete the default contents of the pull request template ... Pull requests that are missing the checklist will be closed." Our draft body for #2391 had ticked both boxes and dropped the nested bullet under the first; the adversarial review caught it against the guide. Its second round then falsified the changelog row the first round had supplied. | A template's default contents stay in the body verbatim, nested bullets included, and they count toward the word limit; read the contributing guide at the head being filed, because the rule can be hours old. A changelog row is outbound text like the body: it names the placeholders a user would search for and claims only cases that differ between base and head, and wording a reviewer supplies is attacked in the next round like any other. |
| 2026-09-13 | google/benchmark #2294, apache/commons-codec #443 | Two maintainers, in one week, asked for comments added to their code to be taken back out, after the density rule had let them through as matching the file. Jeff: "never add comments to someone's code. That's pretty much standard for production level code." | A branch adds no comment to a repository it does not own, whatever the file's own density; a comment already there may be reworded, a new file may open with the header its neighbours carry, and a maintainer's explicit request for documentation is the one exception. comment-check.sh, run by housebroken branch, refuses; housebroken file requires the branch stamp. |
| 2026-09-13 | rouge-ruby/rouge #2332, fastapi/typer #1952, google/benchmark #2294 | Pull requests were filed into a repository whose contributing guide bans LLM use, one whose organization template asks for an AI disclaimer and a discussion first, and one whose AGENTS.md requires AI use to be disclosed in every contribution. The policy reader printed "clean" for all three: its pattern had no plain "AI", its file list was exact-case so Contributing.md was never read, it never read templates or the organization's .github repository's templates, and it exited 0 on a match. |
The policy is read from the repository's own tree listing without regard to case, templates and agent instruction files included, with the plain word in the pattern, and it ends with a reading: BAN, DISCLOSE, MENTION or CLEAN. A project that bans AI-written contributions never gets one; a project that asks for disclosure gets a truthful sentence in the body. housebroken file refuses without a fresh policy printout, refuses a BAN outright, and refuses a silent body under DISCLOSE. When a gate is added, it is run over every open pull request it would have screened. |
| 2026-09-12 | google/benchmark #2294, a reply draft | The draft called the millisecond coefficient "exact, 1.0003662332906062e-03", a value from the local build; all 60 CI jobs print 1.0003662332985102e-03. Three review rounds passed it; a fresh review read the CI job logs. |
A number in outbound text is checked against the maintainer's own output, the CI job logs, before it leaves. Deterministic on one machine is not the same digits on theirs. |
| 2026-09-12 | amazon-ion/ion-java #1165, a reply draft | The draft said "The only behaviour change is still getMillis() before 1582-10-15". The diff also moved calendarValue()'s cutover for every date, so it no longer equals a default GregorianCalendar, and moved forMillis()'s lower bound. Two review rounds passed it. |
"Only" about behaviour is a claim about every public method the diff touches: list them from the diff and probe each against the base before writing it. |
| 2026-09-12 | google/benchmark, apache/commons-codec, amazon-ion/ion-java | Posting steps waited for green CI before a reply. On all three a fork's workflows wait for a maintainer's approval, and housebroken verify printed PR-VERIFY: OK with no checks while they waited. |
A check that cannot fail is not a check: a run awaiting approval is reported as not settled, never as OK, and a reply that answers a question does not wait on CI a maintainer has to start. |
| 2026-09-12 | google/benchmark #2294 | CI was read only at the head, 88 green. The runs a maintainer had approved at the previous head had a red job: complexity_benchmark aborted on windows-2022 when a CPU RMS printed -nan(ind). It was the project's own flake, seen on main twice before, but nobody on our side had looked, and its red X still showed on the commit list. |
Read every workflow run on every commit of the pull request, not only the head's rollup; for each red job, quote the failing line and show from the project's own runs whether it is ours. |
| 2026-09-12 | housebroken 0.5.1, the review and file gates | A fresh review found that file-pr.sh passed gh argument forms it never parsed (-F<file>, --fill, --head, a second -R), so gh would file a body, branch or repository neither gate read; and that review check took the first verdict and any text line anywhere in the report, so an appended DO NOT POST round passed. Both reproduced on the shipped scripts. |
A gate that wraps a command line whitelists the arguments it lets through; a check that binds a report reads only the header it binds. Fixed in 0.6.0. |
| 2026-09-12 | the verification scripts around a review | A reviewer's background test was still running in a fixed work directory half an hour after its completion notice. A validation run of the same script deleted that directory and put the base's file into its replacement; the reviewer then reported a pass "against the rework" for a tree that no longer held it. | A script that works in a copy makes a fresh directory per run (mktemp -d), never a fixed path it deletes first, and an agent's completion notice does not mean its background work has stopped. |
| 2026-09-08 | apache/commons-lang #1783 | "I think this PR creates 2 bugs so it looks like we are missing some tests since the build was green, so please add these missing tests and update the main code ... Static methods do not dispatch to subclass implementations ... Reflection ignores the receiver when invoking that static method." And: "Your AI is imagining things when it talks about PR #1427 because that PR was closed without being merged." Converted to draft. | A change to which reflective member is returned is a semantics change, not a lookup tweak. Test the static grid before filing: a package-private subclass hiding its public parent's static method, a static interface method (never inherited, invoked without a receiver) beside an instance method of the same name, and the JDK case. The body's history claim had been transcribed from the prior-art table, whose closing ruling was "in favor of git master": name the commit git blame gives, never the pull request a table says was closed. Both cases reproduced on the first commit, fixed as a second commit with his own examples as tests, body corrected. The review itself was missed for forty minutes because the check read the commenter list instead of the bodies; the sweep now has an issue for that. |
| 2026-09-08 | apache/commons-lang #1784 | "the main changes look good but I think you need more tests, specifically: Extend the regression tests for subtraction. The new test covers both denominator branches for addition, but its only subtraction example has coprime denominators after reduction. ... Exercise both operands being unreduced and minimum-integer normalization." Then, sixteen seconds after CI completed on the added cases: "looks good, merged" | When a fix adds a step to a helper shared by several public methods, the test is a grid: every public operation through every branch of the helper, a cancellation to zero, and the signed boundary. The first commit had tested addition through both branches and subtraction through one; the reviewer enumerated exactly the empty cells. Four cases added as a second commit, merged the moment CI went green, one day after filing, before any reply was posted. |
| 2026-09-08 | aws/smithy-go #706 and #707 | "Did you encounter this in a real client call? ... If you generate a smithy client today it will use schema-based (de)serialization by default and won't use these APIs." And: "Does the new schema-serde protocol implementation also have this issue? ... this should be effectively dead code now." | Before filing on a library with two generations of an API, find which one shipped clients use and audit that one; a finding on the legacy path is answered with what the new path does. Here the new JSON codec already had the fix and the new CBOR path had the defect with no ceiling at all, which turned a dead-code report into a live one. Answer the reachability question with the file and line, and say plainly when the finding came from an audit rather than a service call. |
| 2026-09-08 | google/brotli #1537, #1538, #1539 | #1538 approved without a word. #1539 approved, then: "Does c/dec/decode.c still need static_init.h include (since invocation is moved to state)? Also, in state invoke ensure before decoder start callback, so that possible initialization cost is not accounted towards decoder instance." #1537: "Comment in CopyStat is a bit confusing (because we enter it if there were no failure) and not necessary (explains not doing something we no longer do here). Please drop it." |
When a change moves a call out of a file, take its include with it; when it moves a call into a timed section, put it before the timer starts. And the comment rule a fourth time: a comment that explains what the code no longer does is confusing by construction. Both reworks pushed as second commits within the hour, replies ten minutes apart. |
| 2026-09-08 | ARM-software/astc-encoder #669 | "Closing - no user-visible bug." after the concession that every loader checks its read and only a transient allocation remained. | The closing line is the whole rule: a change with no user-visible effect is not a pull request, whatever it saves. Conceded in one reply, closed by the maintainer six hours later; nothing further said. |
| 2026-09-08 | uuid-rs/uuid #907 | "This needs some cleanup before it's ready. Particularly around the comments. Most of them should be removed." (3 September) Then, on the reworked commit: "Thanks @lenamonj!" and merged (8 September). | The same rule as console #296, stated by a second maintainer the next day: strip every added comment from code that has none. Reworked as a second commit; merged five days later without another word. |
| 2026-09-08 | ARM-software/astc-encoder #669 and #670 | "Is there any actual bug here? Either a security issue where we read out-of-bounds, or a case where we claimed to succeed but returned incorrect output? If yes, please provide specific reproducer data files." And on the issue: "If an input file is corrupt, allocates lots of memory and then either works with more memory, or fails due to size but safely exits, this doesn't really seem like a problem." | A hardening change needs an out-of-bounds read or a wrong-output reproducer in the body, as a file or a generator, not a peak-memory number. A transient allocation that fails safely is not a bug to this maintainer. One sentence in the #669 body claimed a success with a truncated buffer that could not be reproduced; every loader checks its read. Verify every claim in a body against the fresh clone before filing, and concede the moment one fails. |
| 2026-09-08 | ARM-software/astc-encoder #668 | "Is this behavior actually reachable? Do you have any cases that trigger astcenc to actually try to convert a negative value into fp16? The only case where I think its possible gets clamped to zero before reaching the fp16 path." | A one-token SIMD fix is judged on reachability, not on the intrinsic. The body must name the input that reaches the path and the command that shows the difference: here a half-float DDS with negative texels under -pp-normalize, where AVX2 output differs from the other ISAs on main. Put the reproducer in the body the first time. |
| 2026-09-07 | google/benchmark #2294 | "unnecessary comment" on a one-line comment above the fix; "why are there defaults here?" on two default arguments added so ten existing call sites stayed untouched | Comment density is judged on the lines around the change, not on the file: one comment in a function that has none is one too many. A default argument, overload or wrapper that exists only to leave call sites unchanged hides the change; touch the call sites. |
| 2026-09-07 | cloudflare/circl #699 | "More precisely: ecdh GenerateKey sometimes reads an extra byte on purpose to break exactly what we tried to do here." Then: "Adjust the comment and this is good to merge." | The one comment you do write states the library's intent, not the observed behaviour. "May read an extra byte" reads as not knowing why; "reads an extra byte on purpose, to defeat deterministic derivation" reads as knowing. A review bot had flagged the sampling loop as a timing leak; the reply cited the RFC section and the maintainer wrote "agreed". Answer a bot with the specification, once, and let the maintainer rule. |
| 2026-09-07 | Kotlin/kotlinx-datetime #649 | "The fix looks correct, but the tests can be improved." The test belonged in the JVM-only suite whose helper checks every pattern against java.time. | Put the test where the project's own comparison lives, even at lower platform coverage, when that is where the maintainer proves things. Reworked as a second commit with a three-sentence reply; merged an hour later. |
| 2026-09-07 | Shopify/toxiproxy #770 | Nothing. The CLA action posts no comment and writes its instructions into a failed job's log. | A red check named cla or license with no comment on the thread means read the job log. The contributor agreement card records which organizations' bots are silent. |
| 2026-09-07 | luau-lang/luau #2738 | "Fix for this is already planned for a future Sync." | Some projects land changes through an internal sync. The bug was confirmed and the pull request closed. No public search could have found the fix; the residual risk the prior-art gate cannot remove. |
| 2026-09-06 | webmozarts/assert #366 | "Narrowing the methods would mean a new major release." | A change that rejects an input the project tolerated is breaking on a stable major, however wrong the old behaviour looks. It becomes an issue, never a pull request. |
| 2026-09-04 | thephpleague/csv #591 | "I will close the PR. Thanks for submitting it." after a discussion of whether an unparseable date should reset the field or keep it | The maintainer's preference for the tolerant behaviour stands. Concede in one sentence and let them close it. Same class as assert #366, four days earlier; two closures made the rule. |
| 2026-09-04 | apple/swift-http-types #153 | "What is the use case for a relative URLRequest / HTTPRequest? We had some related discussions in #98" | Closed issue #98 had already ruled. Nobody had read it. The prior-art printout now quotes the closing ruling of every closed item on the touched files. |
| 2026-09-04 | apple/swift-log #503 | Two review rounds: use the package's own Lock instead of NSLock in the compatibility test; then approved, merged three days later. | Use the project's own primitives in tests, not the platform's. Rework as a second commit so the reviewer sees exactly what moved. |
| 2026-09-04 | apache/commons-text #768 | Asked for a negative case inside the valid window and a direct call with explicit lengths. | A maintainer's test request is the review. Reworked and replied within the hour; merged fifty minutes after the reply. |
| 2026-09-03 | apple/swift-log #504 | Asked for the documentation-only form of the fix; the assertion stays. | When the maintainer wants the smaller change, ship the smaller change. Merged the same night. |
| 2026-09-02 | console-rs/console #296 | Comment density and the project's own CI gates the author had not run. | Run every gate in the project's pull request workflow before filing, mutation gates included. No comment in code that has none. The comment census became a script after the third maintainer said the same thing. |
| 2026-08-31 | fastapi/typer #1946 | "Closing, violates user contribution guidelines. ... 1881 is a PR, you don't close a PR with another PR" | The duplicate search had returned pull requests mixed with issues and nobody checked the type. The prior-art printout takes the type from the API field and never infers it. |
- Reworks done as a second commit with a one-sentence reply were merged within the hour four times (commons-text, swift-log #504, kotlinx-datetime #649, circl #700) and within three hours a fifth (circl #699). The maintainers were ready; the pull request had to be in the shape they asked for.
- Every closure that was a mistake on this side was a mistake of reading: a closed issue not read, a search result type not checked, a stable-major contract not respected. Every one became a gate that reads for you.
- Three separate maintainers objected to comments in code that has none
before the census existed, and a fourth did after it existed, because the
census measured the file and the maintainer measured the function. The
gate moved to match the judge.
| 2026-09-09 | cisco/libsrtp #822 | We diagnosed that
configure.acforcedPKG_CONFIG --staticand broke every OpenSSL build on stock Ubuntu, then filed an issue rather than a pull request because dropping the line narrows the generatedlibsrtp3.pc. Another contributor opened #823 seven hours later citing our issue, deleted the line we named, and the maintainer merged theirs. | Step 3 rewritten: a narrowing or default-changing fix now goes as a pull request whose body names exactly what breaks and offers to narrow or close it. The issue route survives only for a policy that demands discussion first or a real compatibility break on a stable major, and even then the patch is pushed to a branch and linked. | | 2026-09-09 | the desk itself | A census tick enumerated 68 of 83 open pull requests and said nothing, because GitHub search returns a short page when throttled; the same day a CI watcher exited on zero pending checks before any existed, and a red-then-green harness reverted with a git command that silently no-opped. All three reported health while seeing nothing. | inbox.sh compares the rows returned against the count the search API reports and prints CENSUS INCOMPLETE with both numbers (0.2.6). The general rule: before arming a watcher or an experiment, ask what it would print if the thing never ran, and require positive evidence rather than absence. | | 2026-09-10 | SethMMorton/natsort #196 | The pull request carried two red checks for hours and the desk never flagged it. inbox.sh read each pull request withgh pr view ... || continue, so a lookup that failed under API throttling skipped that pull request in silence, and an unreadable pull request looked exactly like a clean one. | The census now counts what it could not read and prints CENSUS UNREADABLE with the names. Proven by a test that stubs gh so every lookup fails: it fails on the old code and passes on the new. Same family as the truncated-search guard in 0.2.6, one layer down. | | 2026-09-10 | adobe/elixir-styler, not filed | A patch removed a rewrite the repository documents in docs/general_styles.md with a runnable before and after example, and never touched that file. The code survived five lines of adversarial attack; the deliverable was still incomplete, and the maintainer would have opened it to find his own documentation contradicted. | Before filing a change to documented behaviour, grep the repository's docs, README and CHANGELOG for the behaviour being changed, not just the code. A rewrite rule, a default or a public promise usually has prose somewhere that has to change with it. | | 2026-09-10 | PaloAltoNetworks/docusaurus-openapi-docs #1624 | An adversarial review found the new test file committed 100755 where every other test in the repository is 100644, which GitHub renders in the diff. Fixing the second finding then produced a worse one: the corrected documentation sentence was written into the working tree, the branch was rebuilt withgit reset --soft <base>, and the commits carried the old wrong sentence, because a soft reset leaves the index at the previous commit and the corrected file was never staged. The tree was dirty and nobody looked. |branch-check.sh, run ashousebroken branchat step 9. It refuses a dirty working tree, a file whose mode disagrees with its own siblings, a committed file the repository's.gitignoreexcludes, a working artifact, and a branch that is not on top of its base. Step 4 gained the rule the reset broke: after rebuilding commits, grep the committed blob for the change, never the file on disk. Its own ignore check shipped inert in the first draft, becausegit check-ignoreskips tracked files without--no-index; the sabotage proof in the test caught it before the gate ever ran. | | 2026-09-10 | PaloAltoNetworks/docusaurus-openapi-docs #1624 | The final red-then-green gate printedRED arm: 4 passed.git checkout upstream/mainhad been refused because the tree was dirty, the run continued on the patched tree, and the label said RED while the code said green. Only the sha256 printed beside the label exposed it. | Step 4: swap the file's content fromgit show <base>:<path>rather than switching branches, and assert the two arms' hashes differ before believing either result. A command that is refused reads exactly like a command that worked, so the arm must carry the hash of what it actually ran. Second instance of this shape after the WSL worktree checkout on commons-lang #1783. | | 2026-09-10 | PaloAltoNetworks/docusaurus-openapi-docs #1624 | The patch added a sentence to three documentation files: "clean-api-docsremoves only the files the plugin generated." The reviewer falsified it in one pass: nothing insrc/writessidebar.jsany more and the command still deletes it, and a user file namednotes.api.mdxgoes the same way. The claim was ours, added by the patch, not inherited from the project. | The elixir-styler lesson was to update prose the patch contradicts. This is its mirror: prose the patch adds is a claim the maintainer will test.branch-check.shprints every added line of documentation containing only, exactly, always, never, guaranteed, nothing or everything, and asks for each to be falsified against the code. It warns rather than refuses, because some absolutes are true. | | 2026-09-10 | the desk itself | A one-sentence edit to a README, moving no derived number, was gated on a 342-check validator and a round trip to GitHub's markdown API to prove that a<sub>tag renders. The same file used<sub>fourteen lines above. Jeff: "All this stupid validator shit on ridiculously dumb stuff needs to stop." | Verification is proportional to what the change can break. A prose or markup edit that moves no count marker is read, not validated; the full suite runs when a marker, the scorecard, a receipt, an engine file or a hook moves. Ceremony on trivia reads as diligence from the inside and as noise from the outside, and it buries the verification that matters. | | 2026-09-10 | the repository itself | Version 0.3.0 was bumped, committed and pushed, and the release that publishes it was left for a later decision, so a gate written that afternoon existed in git and nowhere anyone could install it. Jeff: "let's make sure all our updates are being pushed to pypi or it kinda defeats the purpose of updating them." The first parity check written to answer him reported "tagged but never released: none" while reading no tags at all, because the git call ran outside the repository and an empty set differs from nothing. | A bump and its release are one action, not two.tests/test-release-parity.shfails when a published GitHub release has no artifact on PyPI, and when the version in pyproject.toml is on neither. It asserts both lists are non-empty before comparing them, because a comparison against a list that failed to load reports parity for the wrong reason. Proven by sabotage: setting the version to an unpublished one fails the test, restoring it passes, and the file hash is identical either side. Publishing is verified by installing the built wheel and running the shipped binary, never by the index listing, which served a stale page for minutes after the upload succeeded. | | 2026-09-10 | microsoft/GSL #1272 | A red-then-green harness restored the file under test withgit checkout -- <path>in atrap. The header fixes were not committed yet, so the checkout restored from the index, which still held the pull request as filed, and deleted an afternoon of work. The script reported RESTORE FAILED only because it compared hashes afterwards; without that line it would have reported a clean run over a tree it had silently reverted. |git checkout -- <path>is a destructive operation on uncommitted work, never an undo. Commit the change before running any experiment that swaps the file, so the restore target is a commit object rather than a stale index. Where committing first is not possible, copy the file aside and restore from the copy, and compare hashes either side. Step 4's existing rule, that a restore must be verified by hash, is what turned this from silent data loss into a visible failure; keep it. | | 2026-09-10 | microsoft/GSL #1272 | The first full run reported100% tests passed, 15 of 15while the test the maintainer had asked for never executed.tests/CMakeLists.txtsetsCMAKE_CXX_STANDARDfrom its ownGSL_CXX_STANDARDcache variable, default 14, so-DCMAKE_CXX_STANDARD=20on the command line was ignored, and the C++20-guarded test and five new static_asserts were preprocessed away. A green suite proved only that the guards compiled. | A pass count is not evidence a specific test ran. After adding a test, assert it by name in the runner's output (--gtest_filterand a match on the OK line), and read the project's own CMake for the variable that actually selects the standard rather than assuming the CMake built-in. Same family as the sabotage proofs: ask what the run would print if the new test did not exist, and make that answer different from success. | | 2026-09-10 | microsoft/GSL #1272 | A reply draft said "The five new static_asserts fail on main." An adversarial review falsified it two ways and the measurement was reproduced: reverting only the rework fails four of five, because fordyn_array_iterator<const T>the oldconst_referenceandreferenceare the same type, so the defect never affected the const iterator; and against literal upstream/main five fail but so do twenty-one other things, since that state predates the pull request's own first commit. The number came from a red arm that grepped only forstatic assertion failedand so never saw the other errors. | A count taken through a filter describes the filter, not the run. Before a number goes in a body, state what it was counted against, count the total as well as the matches, and name the revision: "against X, N of M". A red arm proves a fix is load-bearing; it does not license a sentence about how much of the failure the fix owns. Cutting the claim was cheaper than qualifying it. | | 2026-09-10 | the door itself | Jeff, on being told an adversarial reviewer had caught an overclaim the author missed: "How is it that Sonnet is better than Opus 5 on Max??? That's like completely wrong imo." The honest answer was not a model ranking. Every gate in this repository reads code: modes, trees, diffs, comment density, red-then-green, CI. Nothing read the prose. Every defect caught in-house that day was caught by a control that printed a hash or a count; every defect that reached a reviewer was a sentence, where there was no control at all. The author had also already concluded the patch was good before writing the sentence, which is the seat, not the model. |claim-check.sh, run ashousebroken claims, listed at step 6 for a body and step 10 for a reply. It lists every counted or absolute claim in outbound text so each is re-measured against the exact revision its sentence names. It deliberately does not judge: the sentence that caused it named its revision and read perfectly. Listing is the gate; the re-derivation is the author's. | | 2026-09-10 | claim-check.sh itself | The first draft could not catch the sentence it was written for. Its pattern wanted the noun beside the number, and "five new static_asserts" has a word in between, soThe five new static_asserts fail on main.passed clean. It had appeared to work only because in the real draft that sentence shared a line with "15/15". | Test a new check against the verbatim artefact that motivated it, in isolation, before believing it. Not a paraphrase, not the file it happened to live in: the exact string, alone in a file. Third instance in one day of a check shipping inert, aftergit check-ignorewithout--no-indexand a release-parity comparison against an empty list. | | 2026-09-10 | PaloAltoNetworks/docusaurus-openapi-docs #1624 | A check appeared in the rollup with an empty conclusion and was reported twice asaction_required, a workflow awaiting maintainer approval because the author is an outside contributor. It is not. That workflow runs on no pull request in that repository at all: four dependabot pull requests from a branch, not a fork, show the same empty state and merged anyway. The real gate was one required approving review, which every merged pull request there carries and none of ours did. | An empty or unfamiliar CI state is a question, not a diagnosis. Before explaining it by something about your own pull request, look at pull requests that are not yours, ideally ones that merged, and see whether they show the same thing. A cause that is also present on every successful case is not the cause. The same one-line test settled the stylex Vercel failures the same day, by comparing fork against branch pull requests; it was simply not applied here. | | 2026-09-10 | microsoft/GSL #1272 | Copilot's suppressed comment on the first commit: "The comment about testing transitive inclusion is now inaccurate: this file explicitly includes , but the comment says is not included directly. Please update the comment (or remove it) so it matches the current include list and test intent." The fix commit rewrote the comment. The comment was the specification: the file omits<algorithm>on purpose to prove<gsl/dyn_array>supplies it, and the first commit had added the include without needing it, sincestd::sortalready arrives through the header. Caught in the Fable review before the rework went out; the include is gone again and the original comment stands. | When a reviewer or a bot reports that a comment and the code disagree, decide which one is the specification before touching either. A comment that states a test's intent is the spec, and the code is what drifted. "Update the comment" is the bot's suggestion, not a diagnosis. | | 2026-09-10 | NVIDIA/go-nvml #207 | NVIDIA's review bot, four days after filing: "The commit signs off aslenamonj, while the contribution rules require a real name. Amend the sign-off before merging." CONTRIBUTING.md says it verbatim: "You must use your real name (sorry, no pseudonyms or anonymous contributions)." The commit also carriedCo-Authored-By: Claude, a day after the rule against tool trailers was written down. Both came from the machine: the globaluser.nameis a username and the harness appends the trailer, so every commit inherits both unless something refuses it. The DCO check stayed green throughout: the sign-off's email matched the author's, and nothing checks whether the name is real. | Read the agreement's identity clause before the first commit to an organization, and set author, committer and sign-off to the name it asks for; a green DCO check proves the sign-off is present and consistent, not that it names a person.housebroken branchrefuses username sign-offs and tool trailers from 0.5.0; a rule that lives only in prose was inert for the five days in between, and the refusal itself, written on 2026-09-10, sat uncommitted until 2026-09-12. | | 2026-09-10 | NVIDIA/go-nvml #207 | Review before the fix went out found two sentences in the body and commit message that the code contradicts. "A heap overflow ... on everyPath()call for a library opened by soname":Path()caches its result indl.pathbeforedlinfois reached a second time, so only the first call after eachOpenoverflows. "Path()on main segfaults inside cgo": measured three times at the stated 224-character directory,Path()returns the right string and the SIGSEGV arrives in the followingdlclose. | A frequency claim about a bug ("every call", "always") is checked against the caches and early returns in the same function before it is written. A crash claim names the frame that crashed, from the trace, not the function that holds the bug: heap corruption surfaces at a later allocator or loader call. | | 2026-09-10 | the review of go-nvml #207 itself | The first red-then-green run for the long-directory crash reported PASS on both arms. The script had restricted PATH,ldconfigwas not on it, the library copy never reached the long directory, and both arms ran against the system library. The only sign was onecp: cannot statline in the middle of the output. | The arm must print evidence that its precondition held: assert the copied file exists, and print the path the code under test actually resolved. The rerun printed a 232-character path on both arms, then SIGSEGV on main and a clean exit patched. Fourth instance of an experiment that reported health while seeing nothing. | | 2026-09-10 | microsoft/GSL #1272 | The new#include <span>was guarded on__cpp_lib_span, a macro that only<span>and<version>define. It compiled on gcc because libstdc++'s<ranges>happens to include<span>; on a standard library where it does not, the test the maintainer asked for would have been compiled out in silence or the build would have failed. The repository's ownspan_tests.cppguards the same include on__cplusplus >= 202002L, with MSVC built under-Zc:__cplusplus, and that guard has passed the whole CI matrix for years. | A feature-test macro can gate an include only if something before the gate defines it. Before writing a guard, grep the repository for the same include and copy the guard it already uses; a guard that works on the local toolchain by transitive luck is a red check on someone else's. | | 2026-09-12 | microsoft/GSL #1272, a reply packet | The packet for a reply to a Copilot review argued that the header "already does this arithmetic on the same null pointer" and citedoperator pointer()and_Unchecked_end(), both found by grepping for_ptr +anddata() +. Both sit under#ifdef _MSC_VER, so neither gcc nor clang, the two compilers the evidence ran on, ever compiles them. An adversarial review passed the packet; a re-read of the lines around the grep hits caught it. Nothing had left the machine, and the reply itself did not cite it. | A grep hit is a line without its context. Before citing code as precedent, read the enclosing preprocessor guards and scope: a line under a platform guard is precedent only on that platform. The claim-list gate reads outbound prose; nothing reads the context of a quoted line. | | 2026-09-12 | the same packet | Three numbers in its first draft were written from memory of a JSON dump rather than recounted: "all 90 checks green" (88 passed, 4 skipped), "every issue (10) closed since 09-04" (4 of the 10 were this repository's own; 6 outside), and a coefficient range whose upper end came from the run with the values scaled by 1000. All three were caught by recounting before the review ran. | Every number in a document that goes to a reviewer is recounted by a command at the moment it is written, and the count names what it was taken over. Same family as the 09-10 row on counts taken through a filter, one step earlier: a count taken through memory. | | 2026-09-12 | this repository's own uncommitted work | Asked what the uncommitted changes to branch-check.sh were, I reported that both files had been rewritten with CRLF line endings, 170 and 99 lines carrying a carriage return, and that bash would choke on them. Under Git Bash on Windows,grep -c $'\r'matches every line: an untouched control file scored 149 of 149, and a byte count found no carriage return in any of the three files. Two rewrites that changed nothing were the tell, and they were read as a stubborn file rather than a lying count. The real damage was elsewhere: aprintf '%s\n'split across two lines and a lost line continuation, both backslashes eaten on the way into the file. | Count bytes with a tool that reads bytes, and run the same check on a file known to be clean before believing a number about a suspect one. When a fix leaves a count exactly unchanged, suspect the count before the fix. | | 2026-09-12 | the review of google/benchmark #2294 | The adversarial reviewer copied the clone withcp -aso its experiments could not touch the original, then rebuilt withcmake --build. CMake's generated makefiles hold absolute paths to the tree they were configured in, so every arm compiled the untouched original, printed "Built target" for everything, and reported results for code it never built. It caught this only because it required each arm's build log to show a "Building CXX object" line for the file that arm changed. | A copied CMake build directory builds the directory it was configured from. After copying a clone, reconfigure (cmake -S . -B build) before the first arm, and require each arm's build log to name the file the arm changed. Fifth instance of an experiment reporting health while seeing nothing. | | 2026-09-12 | the door itself | Step 9'shousebroken branchanswered "unknown subcommand", and there was nohousebroken claims: the command on the Git Bash PATH, ~/.housebroken/bin/housebroken, reports "(unversioned)" while this repository and PyPI are at 0.4.0. Both gates were run by hand from scripts/. The rows above that say "run ashousebroken branch" were true of the repository and false of the machine. The first draft of this row said v0.3.0 was the newest tag, read from a clone that had not fetched v0.4.0. | A gate exists where the installed command runs it. After a release, install it on every host that walks the door and checkhousebroken versionagainst pyproject.toml; install.sh now writes the version beside the dispatcher so a script install can answer. A session that meets an unknown subcommand upgrades first rather than quietly falling back to the script. | | 2026-09-12 | this repository's README | The badge under the title said "shellcheck clean". Preparing 0.5.0, shellcheck found an unused variable in scripts/inbox.sh and two findings in tests/test-claim-check.sh, and the same three were already there at 0.4.0, so the badge had been false for at least one release. Nothing checked it. | A badge is a claim with no sentence around it, so the claim listing never sees one. Re-derive every badge before a release. The three findings are fixed in 0.5.1. | | 2026-09-12 | google/benchmark #2294, a reply draft | The draft saidcpu_coefficientruns "3.0e-10 to 3.5e-10 ms per N across runs here". The range came from three runs. The adversarial review ran ten: 2.91e-10 to 3.97e-10, six of them outside the stated range. | A range quoted in outbound text is a claim about every run a maintainer will make, not about the runs you happened to make. Quote an order of magnitude, or run enough to bound it, or say only that it varies. | | 2026-09-12 | microsoft/GSL #1272, a reply draft | The draft said gcc and clang "rejectnullptr + 1in its place". What was tested was aconstexpr int*null plus 1 in constant evaluation. A literalnullptr + 1also fails, but only becausestd::nullptr_thas no+, so a maintainer who tried the sentence as written would see the right result for the wrong reason. | Describe the test that ran, in its own terms. A sentence that paraphrases evidence can stay true while pointing at a different experiment. | | 2026-09-15 | IBM/sarama #3740 | The fix bounded the subscriber search with aprobecounter beside the existing cursori, then added the counter back into the cursor. The approving reviewer, puellanivis: "We could probably do this by reworking the existing control flow instead", and gave the loop with the cursor itself bounded byfor range n. Same member chosen, same cursor advance, three fewer moving parts. The adversarial review had passed the patch, because nothing in its brief asked whether the code was in the form the best engineer in that language would write; it attacked claims, proof, CI, scope, prior art and prose. | The brief gains attack item 4, the form: the reviewer writes the changed block as the strongest engineer in that language would, in the repository's idioms and at its manifest's language version, and a smaller or clearer version is a SHOULD FIX with the rewrite in full. A rework request is a review round spent on style instead of the merge. | | 2026-09-16 | teslamotors/vehicle-command #479 | The maintainer, sethterashima, on the helperisRequestTooLarge: "Useerrors.Is(....)instead of a helper function." The elided argument does not exist:*http.MaxBytesErrorcarries no sentinel and noIsmethod, soerrors.Is(err, &http.MaxBytesError{})is false for every read past the limit. The literal ask was impossible; the ask underneath it, no helper, was not. | A reply to an impossible ask leads with the ask done (the helper is gone, the check is inline) and gives the one reason the literal form fails in a clause the maintainer can paste,errors.Is(err, &http.MaxBytesError{})is always false. It never opens by correcting him. | | 2026-09-16 | teslamotors/vehicle-command #479 | The adversarial review of the rework found the PR's own test, unchanged since filing on 6 September, posted to a route the proxy does not have, covered one of the three sites, and on the red arm sent a real POST with a forged bearer token tofleet-api.prd.na.vn.cloud.tesla.com, because it was the only test in the file that reached the forwarder without stubbingp.client. Nothing in the filing gate ran the red arm with egress blocked, so the test that proved the defect also proved it against production. | Step 4 gains a rule: the red arm runs with egress blocked (HTTPS_PROXYandHTTP_PROXYat a black-hole address, or the equivalent) and a test that then names a host has left the machine and is rewritten to stub the client the way its neighbours do. Filed as an issue for a mechanical check. | | 2026-09-16 | teslamotors/vehicle-command #479 | The rework was first pushed as a second commit, per step 10. The reviewer checked the repository's merge style: no merge commits since 2024 and PR #459 sitting in main as its own two commits, so this repository rebase-merges and a second commit would land in main permanently, carrying comment deletions and a rewritten test under a one-line subject. | Step 10's "second commit" is conditional on the merge style: whengit log --mergeson the default branch is empty and prior PRs sit in main as their own commits, the rework is squashed into the original commit with its message amended, and the maintainer reads the force-push compare instead. | | 2026-09-17 | NVIDIA/go-nvml #207 | Six days after filing with a DCO sign-off, the maintainer, tariq1890: "can you take care of this too?" The same maintainer had asked the same on #178. CONTRIBUTING.md names only the DCO; the Verified requirement lives in the maintainer's habit, and the repository's merged outside commits all show it. No signing key existed on the account, so the answer took a key, a scope refresh and a force-push. | Step 7 gains a read: before the first pull request to an organization, check whether its recently merged outside commits show Verified, and when they do sign from the first commit. A signature added later costs a force-push, a re-approval of CI and a round. | | 2026-09-17 | NVIDIA/go-nvml #207, the reply | The adversarial reviewer's round one made the reply tell the maintainers their CI "needs /ok-to-test ", inferred from a 45-second gap between one comment and one run. Round two read the run objects: every run is created at push time and started by a maintainer in the Actions UI, and on #204 the run started fifteen seconds before the comment. The sentence was false and would have been posted in the imperative to the people who own the workflow. | A claim about how a project's CI is triggered is read from its workflow files and from the run objects (run_started_at,triggering_actor,run_attempt), never inferred from comment timing. The brief's item 3 names those fields. A reviewer's finding is evidence only when the evidence is in it. | | 2026-09-17 | NVIDIA/go-nvml #206 | A signing force-push was due on a pull request filed four days earlier, before the no-added-comment gate existed. The patch carried a four-line paragraph under a doc comment in a file where every function has exactly one comment line, and four comment lines in a test in a package where four of five hand-written tests have none. Left in, they would have gone up again under a fresh Verified head. | Any force-push re-runs the current gates over the whole patch, not only the part being reworked: a head that is rewritten anyway is the cheapest moment to bring an older patch up to the rules that came after it. The reply names what was dropped so "nothing else changed" is checkable. | | 2026-09-17 | NVIDIA/k8s-device-plugin #2002 | Eleven days after filing, the maintainer, abrarshivani, on the new test file: "Can we cover the config default behavior here? In particular, it will be useful to verify that an omitteddeviceListStrategygets theenvvardefault, while an explicitly empty list does is rejected." The patch had proved the rejection by unmarshalling JSON into the struct and calling the constructor; the default that must keep working was asserted nowhere. | A rejection test ships with its complement. When a change refuses an input, the test also drives the input that must keep working through the same production entry point the refusal goes through (hereNewConfigwith the flag at its default), so the maintainer sees both halves of the contract in one table. A test that only unmarshals proves the parser, not the behaviour. | | 2026-09-17 | NVIDIA/k8s-device-plugin #2002 | The same review, on the constructor guard: "Can you also updateAllCDIEnabled()to returnfalsewhen no CDI strategy is enabled? This avoids the case whereAllCDIEnabled()returnstruewhen the strategy list is empty, which also contributed to this issue, and keeps the behavior consistent." The patch had guarded the constructor and left the predicate that was vacuously true over the empty set untouched. | A guard at the constructor is half an invariant. Every function whose result over the empty input let the bug through gets fixed in the same pull request, and the rework proves the non-empty results unchanged (an exhaustive run over all sixteen enable-subsets, differing only on the empty set) so the maintainer can accept it without re-deriving the predicate. | | 2026-09-27 | google/osv-scanner #3107, the reply | G-Rath asked: "Do you have a reproduction with real world advisories using the CLI or library?" The reply said "Yes" and pasted four real LSN advisories fed togrouper.Group, an internal package. His answer: "That's not a reproduction though that uses the public API". The reply had been rated high confidence: the evidence was solid, but it met one of the two conditions the question named and claimed both. | A maintainer's question is decomposed into every condition it names before a reply is drafted, each condition gets the evidence the reply carries for it written under it, and a condition with no evidence makes the reply not ready. "Reproduction" means the shipped command line or an exported function of a public package on a fresh clone, with its real output pasted. The Opus brief for a reply carries the decomposed question so the reviewer attacks the match between question and answer, not only the facts. | | 2026-10-05 | NVIDIA/go-nvml #206 | tariq1890: "Please squash the commit history." The PR had grown a second commit when his one-line SPDX header suggestion was answered as a separate signed commit on 09-21. | A maintainer suggestion applied before any approval is squashed into the commit it corrects, not stacked on it; a second commit is for rework the reviewer must diff, and a header line is not that. Squash on the same base so review anchors hold, and say so in the reply. | | 2026-10-06 | ARM-software/astc-encoder #668 | Merged by solidpixel 28 days after filing with no review comment; his only question had been "Is this behavior actually reachable?" on 09-08, answered the same morning with the input path, the flags that expose it and a reproducer generator, which he thumbs-upped and then left for four weeks. #669 and #670, filed the same night with measurement-only evidence, were closed within a day. | A reachability answer that names the input, the command and the differing outputs is the whole case for a SIMD fix; after it, silence is a queue, not a verdict. Never nudge. | | 2026-10-06 | aws/smithy-go #706 | lucix-aws, 09-08: "If you generate a smithy client today it will use schema-based (de)serialization by default and won't use these APIs ... the short-term intent is that these old encoding/json and encoding/xml Value APIs become deprecated." Closed by us four weeks later. | A correct fix to an API the maintainer is retiring is not a finding worth a pull request. Before filing, read the latest release notes and open deprecation threads for the package touched; a path on its way out gets an issue at most, or nothing. | | 2026-10-06 | NVIDIA/go-nvml #206 | tariq1890 approved 37 minutes after the squash he asked for and merged the next morning, 30 days after filing; the Devin bot race finding had been answered with the lock change three weeks earlier and the thread then sat until the squash request. | At NVIDIA the review bot and the human review are separate queues: the bot finding is answered in code, and the human engages when the PR is one clean signed commit. A mechanical ask answered within the hour is what moves the human queue. | | 2026-10-07 | Shopify/toxiproxy #770 | iyaz-shaikh: "This was handled as part of #774, so I'm opting to close this PR." His #774, filed 25 days after ours and cross-referencing it, carried his own chunk fix, 400 validation and an API auth feature, and merged the day he closed ours. | Some maintainers fix a reported crash in their own pull request rather than merge an outsider's, and a filed pull request is then a bug report with a patch attached. The finding still landed; it is a fixed-upstream bullet, never a merge, and the close needs no reply. |