Add Adversarial Execution Evidence predicate (v0.7) - #570
Conversation
Evidence produced by executing an untrusted artifact against a digest-committed adversarial corpus inside a containment substrate: recomputable outcome, attack-granular coverage binding, independently signed intercept records under a predicate-level batch root, and explicit negative scope. Verdict semantics deliberately excluded; they belong in a downstream summary predicate such as VSA. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
|
Taking you up on the early-draft invitation. Mostly field-level, plus one gap I think this predicate is unusually exposed to.
That interacts badly with the property I like most in the draft. Because This is not the coverage question from #557. You have already drawn the distinction once.
On your two questions, though you asked the maintainers and I would defer to them on the first. Verdict-free evidence reads as its own predicate rather than SCAI, since the recompute rule is normative here while SCAI leaves attribute and evidence formats to producer and consumer with no re-derivation requirement. On I can send the three-value vocabulary as a suggested diff if you want it concrete. |
|
You found the real hole. The recompute is the property I care most about, and you're right that it's exactly what makes the vantage gap dangerous. A fail derived from rows the artifact itself reported verifies and recomputes identically to one the substrate intercepted. And because the recompute is what buys confidence, everything downstream reads it as checked. I've adopted your fix as proposed and pushed the revision, so what is below is the committed text itself. The actualLayer clean-run behavior is pinned too. On a clean row the producer MUST emit the literal none, explicit rather than omitted, so that "nothing needed to act" is distinguishable from an accidental omission. Your positioned-to-see question turned out to be the best argument for basis on clean rows as well: a clean row carrying substrate_observed states the substrate had vantage and observed nothing, which is the claim a pass actually rests on. That keeps actualLayer answering only which layer acted, and keeps vantage from being loaded onto it. On the rename, agreed and done. Two accepted spellings means two canonicalizations for the same content, so no alias; the old spelling is rejected, and migrating old producer output is our producer's problem rather than the wire format's. The field is On SCAI, same read for the same reason: the normative recompute rule is what makes this its own predicate, since SCAI deliberately leaves attribute and evidence formats to producer and consumer with no re-derivation requirement. The call stays with the maintainers, but a second reviewer landing on own-predicate off the recompute rule is the strongest version of the argument I could bring to vetting. The suggested diff is the smaller half of your offer, and I'd take the stronger version. The vocabulary is committed text now, and what I want from you is the read you gave the recompute, aimed at the observation that fits none of the three cleanly. The seam I'd probe is between substrate_observed and inferred: a substrate-adjacent monitor with real vantage that isn't the enforcing substrate itself. If that deserves a fourth value, now is the time, while there are still no deployed readers to keep. Edited 2026-08-31: prose only, for readability. No field name, vocabulary value, version or claim changed. |
|
Read it the way you asked, against the recompute rather than the prose, and the three values hold up better than the seam you flagged. Let me take that seam first, then hand you a sharper one. The substrate-adjacent monitor does not want a fourth value, because what separates it from the enforcing substrate is enforcement role, not basis. A passive tap that saw an egress at the network boundary and the inline gate that intercepted the same egress make the identical claim on that row: each observed the event at its own vantage, independent of whether the artifact cooperated. What differs is that one was positioned to act and the other only to watch, which is an enforcement-role fact rather than a basis one, and whether anything acted is the separate question The seam I would actually probe runs inside a single value rather than between two, and it passes through one of your own That matters for the recompute for the same reason unknown basis did. You refused to let an unknown vantage launder down into Everything else in the revision reads clean. |
|
Both halves taken. Your vantage class reading is now the definition itself: in the new text I did not take your cheap fix, for a structural reason rather than taste. A flat "inferred reads no stronger than artifact_reported" binds only the producer who volunteered the weaker label, and with my examples ambiguous the same snapshot diff pipeline could defensibly label substrate_observed, so the line would have penalized the honest producer without constraining anyone else. Your strong version is the one that holds: the independence of an inference's inputs has to travel. The revision at def32f0 does that literally. The directness axis is my extension. It has to live on the wire rather than in producer vocabulary for the same reason basis did: the recompute and the documented gating on the pass side read it, and transient tolerance cannot be gated fail-closed across producers on open per-producer labels. Your transient example also bites hardest where you did not press: on pass. The ordering is two-sided now. While basis bounds a fail over a defined supporting set, method bounds a pass, with (substrate, intercepted) clean rows as the strongest absence claim the predicate can carry and any reconstructed clean row read as tolerating transients between the observed states. Rows fail-closed on either field sit at the bottom of both orderings, which is where a row with unknown vantage belongs. One more change completes your own no-fourth-value argument. You delegated enforcement role to Mechanics: the three 0.4 spellings are rejected with no alias, same protocol as the does_not_assert rename, licensed by zero deployed readers. I resisted a third axis (confidence, attribution) as territory for producer vocabulary. Naming is open. If the maintainers prefer observationMethod or directness over method, that is a cheap rename before vetting and their call. The proto lands once the field shape settles, as the PR body says. The main thing I want your read on is whether the weakest-input rule and the artifact-sourcing criterion close the upward leak you named. Beyond that, check me on where the method/attribution line sits, and on whether a caught-row none covers your passive tap the way you meant it. Yes to second reviewer at vetting. Edited 2026-08-31: formatting only. A few repeated field names lost their backticks after their first mention; no name, value, commit or claim changed. |
92424ac to
4133b14
Compare
|
That is the right cut, and the reason the hunt failed is as you put it: three values were carrying two questions, so no single value could fit an observation that varied on both. Weakest-input over vantage and directness is the composition that makes the two axes independent again, and the artifact-sourcing criterion, a channel the artifact can populate without performing the claimed event, is the sharp version of what I was reaching for with testimony. Egress-capture-is-substrate-because-the-packet-had-to-be-sent is the case that proves the criterion carries its weight. On your main question, the weakest-input rule closes the leak I named, but only for a producer telling the truth about its inputs, and that deserves being exact about because it is the same shape as the gap I opened first. The close is to make the strong label cost the one thing a self-reporter does not have. A row may carry On the passive tap, yes, and making On where the Yes to second reviewer at vetting. On naming, |
…nd pin actualLayer clean-run behavior Add a required basis field to every attackResults row with a closed three-value vocabulary (substrate_observed, artifact_reported, inferred). The recompute reads the row, so the row must carry its own vantage: without it a substrate-intercepted observation and one transcribed from the artifact's own output recompute identically. Missing or unknown basis values fail closed exactly like out-of-vocabulary containmentObserved labels, and consumers can gate on basis; a fail supported only by artifact_reported rows is a weaker claim a consumer may reject. Pin actualLayer on clean rows: producers must emit the literal none rather than omitting the field, so an unacted row is distinguishable from an accidental omission; vantage on clean rows is carried by basis. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
…NotAssert and bump to 0.4 Rename the negative-scope field to the lowerCamelCase convention with a single canonical spelling and no alias: two accepted spellings would mean two canonicalizations for the same content, against the I-JSON discipline the spec holds everywhere else. Versions 0.1 to 0.2 were internal, so there are no deployed readers to keep; migrating old producer output is a producer concern. Bump the predicate to 0.4 and record the review-driven changes (basis field, actualLayer clean-run pin, the rename) in the changelog. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
…o vocabulary labels The actualLayer clean-run rule defined a clean row as any row whose containmentObserved label is not in the caught set. An out-of-vocabulary label is also not in the caught set, and such a row is fail-closed, not clean, so the definition contradicted the fail-closed rule. Restrict the clean-row definition to labels from the published vocabulary. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
…ing to the changelog Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
…irectness axes and bump to 0.5 Review of 0.4 showed the three-value basis vocabulary conflated two orthogonal properties: whether the artifact could forge the observation (vantage) and whether the event was captured as it occurred or derived from state after the fact (directness). The spec's own examples instantiated the collision: a hypervisor snapshot diff fit both substrate_observed (VM introspection) and inferred (post-hoc state diff). Replace basis with two required, closed, fail-closed row fields: basis (substrate | artifact), the vantage of the claim's weakest input with a stated artifact-sourcing criterion, and method (intercepted | reconstructed), the claim's directness with the same weakest-input composition rule. Delete inferred with no successor; reject the 0.4 spellings with no alias, keeping a single canonicalization per content. Make actualLayer required on every row and extend the literal none to caught rows, where it states observed-but-not-enforced, so enforcement role travels there and never leaks into basis. Add two-sided consumer strength orderings (basis bounds a fail, method bounds a pass), a coherence check of row claims against the pinned observation environment, a row-internal check that intercepted caught rows reference verifiable intercept records, and state the row-travel design invariant under Parsing Rules. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
4133b14 to
def32f0
Compare
… coverage records and bump to 0.6 Review of 0.5 showed a substrate basis was a bare producer claim: a row could declare the strongest vantage while referencing nothing, so a self-certified strongest pass or fail was indistinguishable from an earned one. Back every substrate row with substrate-signed coverage at two separate gates. The first gate is validity, and it is a pure function of carried bytes. References must be non-empty, resolve in range, and class-match; every covering payload must be canonical +json carrying the reserved members, with a run binding equal to the one derived from the statement; a row's method may be no stronger than the weakest method its covering records signed; and batchRoot must recompute. Failing any of these makes the attestation invalid for every consumer, and evaluating them is a normative consumption precondition rather than an optional lint, so a consumer that reads result or credits a row must run them first. The second gate is a derived per-row evidence tier holding the one genuinely trust-relative question: whether the covering signatures verify against a key the consumer's policy names as a substrate observation key. Rows resolve to attested, unattested, or declared, a consumer with no pinned substrate root treats every substrate row as unattested, and an unattested substrate row ranks where an artifact row ranks. The tier never alters result, and the may-reject-never-downgrade rule is retained. Solve the clean-row case, where an armed vantage that captured nothing produces no per-event record to sign, with a run-level instrument pair: an arming record stating a live capture vantage was armed before corpus injection, and a sealed record stating the vantage stayed armed to run-end with a zero or self-bounded run-wide drop count and an unchanged posture digest. A clean intercepted row is valid only when both cover it, which closes arm-then-drop, drop-count overflow, and mid-run posture flips in byte-checkable text. Carry the observation vocabulary in the attestation, as labels, a caught subset, and a canonical digest, so the recompute and the validity gate no longer depend on a document that does not travel with the statement and archived attestations stay verifiable. Rename the run-start digest to runEntropy, state its pre-image, and fold it into a versioned run binding so identical-configuration re-runs derive distinct bindings; the binding is anti-splice, not a freshness challenge. Rename interceptRecords to observationRecords and interceptRefs to observationRefs, pin batchRoot to RFC 6962 with domain separation and duplicate rejection, replace the trust-boundary paragraph with a field partition and an honest single-root key model, and state the composition and run-population non-claims. Breaking, with no aliases for the old spellings, under the same single-canonicalization rule as earlier renames. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
…row ordering redundancy The observation-environment list carried two conjunctions, and the clean-row ordering stated self-reported absence twice. Fold the unattested substrate clean row into the single weakest-case sentence alongside artifact clean rows. No normative change. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
|
Both halves of the gate taken. I want to lead with the two places my first pass would have overclaimed, because you would have caught them on the first read and I would rather retire them myself. The revision at b5acaa5 is v0.6; as before, this reply lands after the push so you are reading committed text. The first place is the one you care most about, the recompute. My instinct was to add the coverage gate as a second object every consumer derives beside The second place is the single-key topology, and here I have to concede the frame before I describe the mechanism, because your read falsifies the confident version in two lines. In the deployment we actually ship, the substrate signer and the assembly signer are the same key held by the same operator. Under that key the tier does not defeat a pipeline with no substrate in the loop. An operator holding the key can hand-author an arming record with the right derived run binding and a valid signature without any substrate ever running, and every clean row it points at derives attested. A signature proves key possession, and nothing about it proves a substrate executed. So the named attacker, the self-certified strongest pass and strongest fail, is closed only where the substrate observation key is held apart from the assembly plane, which is a SHOULD we do not yet ship. What the tier buys under one key is smaller and real. Against a party that does not hold a substrate key, a downstream tamperer, it still binds every record to this run so a foreign-run record cannot be spliced in, still commits the whole record set under Now the mechanism, with those two retractions in view. Your rule keys on a substrate-signed record covering the observation, and your pass-side sentence wants intercepted on the pass side resting on the substrate too. Each collides with a line I had already committed, that a clean row's interception has nothing to sign. So the pass side needed the instrument your sentence presupposes without naming. It also needed a second record, because a single arming record proves the vantage was armed at t0 and nothing more, and a vantage dropped one tick after arming would satisfy the strongest pass claim. That is your own "a missing signal is not a clean one" reintroduced on the pass side. So in v0.6 Making "covers that observation" verifier-checkable meant closing three gaps I had left open in the first pass. The record content is producer-defined, so I had to say what a verifier is allowed to assume it can parse: any record covering a substrate row MUST be a canonical RFC 8785 plus RFC 7493 object whose media type ends in +json, with the reserved members at top level, else it covers nothing. Without that a producer with duplicate On the reserved members themselves, I would still call them binding structure rather than observation semantics, and I have kept the substrate's own vocabulary out of the wire. It remains the design's most novel joint and I would rather defend it with you now than at vetting: no vetted predicate reaches into a producer-defined payload with reserved members. The alternative I rejected was a producer-declared pointer map telling the verifier where to look and how to decode, because it hands the join back to the party the gate is aimed at. I also dropped the hard exclusivity rule I floated, that an interception index binds one row, because you would have built the counterexample yourself. One captured TLS flow can genuinely carry two exfil payloads for two attacks, and forcing two records over identical bytes manufactures ambiguity. Instead a shared index is allowed when the committed payload evidences each attack, and a row may carry an optional selector member naming the sub-observation it rests on, parallel to its references, with the token content staying producer vocabulary that nothing normative reads. The anti-double-attribution work is carried by the run binding, the On As an I-JSON reviewer you would have flagged two smaller completeness gaps, and both are fixed. The run binding pins The passive-tap reading you traced is untouched, and the frozen shape is untouched: two required closed axes, weakest-input on each, two-sided ordering. The tier is a derived object like Mechanics: version and Type URI move to v0.6 under the meaning-change rule, which the framework's 0.x versioning licenses; the renames are rejections, not aliases. Proto still waits for the shape to settle. Before the freeze I would point you at three seams. The reserved members, per the flag above, are the load-bearing joint; an adversarial producer laundering through them, or a vetted-precedent objection I have underweighted, is where I most expect to be wrong. The Edited 2026-08-31: prose only, for readability. No claim, figure, field name or version changed; a few repeat field mentions lost their backticks and nothing else. |
Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
The source cites spec:NNN line numbers against a specification file that was not present in the repo (only the JSON schema shipped), so a public reader following a reference hit a dead path. Vendor a byte-verbatim copy of the v0.6 predicate spec (tracking in-toto/attestation#570 at b5acaa5) so the repo is self-contained and the line references resolve. A spec/README records the pin and that the in-toto catalog namespace, not this repo, is the canonical authority. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
|
I have been through v0.6 in full at 484bbe0. The two retractions are the right calls and I will not relitigate them; the four byte-pure validity steps as consumption preconditions, with signature verification as the one trust-relative tier, is the correct partition and it closes the result-only consumer cleanly. Taking your three seams in order, leading with the one where you most expect to be wrong, because I think you are half right about it. Reserved members. The vetting objection has a better answer than the text currently gives itself, and the joint has one real crack. The precedent first. "No vetted predicate reaches into a producer-defined payload with reserved members" undersells your own lineage. RFC 7519 does exactly this: a JWT claims set is producer-defined, yet verifiers read registered claim names ( The crack is in the canonicality gate, and it is an adversarial-producer path. A covering payload MUST be canonical per RFC 8785, and a verifier checks that by re-serializing the parsed object and comparing bytes, or by an equivalent sortedness walk. JCS sorts members by UTF-16 code units; most naive re-serializers sort by code points. Those orders diverge when a member name outside the BMP meets one in U+E000 through U+FFFF: a surrogate-led name sorts first under JCS and last under code points. A producer can mint a payload whose bytes are canonical under one reading and not the other, and the two verifiers then split on "covers" versus "covers nothing", which under your validity gate is attestation-valid versus attestation-invalid on identical bytes. You already close the number half of this with the RFC 7493 safe-integer profile, there so that every rail derives identical bytes; the same clause wants a string half. Cheapest fix: require member names in covering payloads to be BMP-only (ASCII would also do, and your reserved members already are), at which point UTF-16 and code-point order coincide and the divergence is unconstructible. The alternative, mandating a true UTF-16-sorting re-serializer in every verifier, puts the burden on the many rather than the one. runEntropy. I cannot give you challenge-free replay exclusion without state, and I do not believe anyone can: stateless deduplication against a global set is a global view, which is what logs are for. What I can offer is an upgrade to the bound you already have. Your text says the pre-image is the substrate's run-start checkpoint "or beacon head", and I would promote that aside to a SHOULD: make the pre-image include a publicly datable value, a drand round or an epoch ID in the RFC 9334 section 10.3 sense, with the round reference recoverable by the consumer (carrying it in the arming payload as producer vocabulary suffices, since the digest binds it). That buys two things at no new wire members. The arming record gains a proven floor, since a signature over a beacon value cannot predate the beacon round; Run population. You are right, and there is a strengthening that does not pretend otherwise. No self-contained attestation can prove the absence of sibling runs; that is the split-view problem, and fork consistency is the known ceiling for it. But your own design already holds the answer one level down: the sealed record's run-wide drop count and the checkpoint chain's each-interception-carries-a-higher-sequence rule make observation loss gap-evident inside a run, and the identical construction lifts to runs. One reserved pair inside the arming record's signed payload, a monotonic run sequence number and the previous run's binding digest under the same substrate key, makes cherry-picking gap-evident across whatever does get published: a skipped run is a numeric gap, a suppressed-and-rerun is a fork, two attestations sharing a predecessor. TUF's snapshot role is the vetted precedent for signing a population, and SCITT registration (RFC 9943, with COSE receipts per RFC 9942, both published last month) is the consumer-side completion that turns gap-evidence into third-party auditability. So the predicate's line that this is "a claim about the run this attestation carries, never about a run population" stays true, while cherry-picking moves from "outside by design" to "detectable across the published set", which I think is the most a single-attestation format can honestly buy. On One offer, since the validity gate is now four pure functions of carried bytes: when you mint conformance vectors for v0.6, I will write the validity-gate checker from the spec text alone, same discipline as the testigo cross, and we will find out whether these sections determine a unique implementation the way we have been arguing they should. |
RFC 8785 sorts object members by UTF-16 code unit, but a verifier that compares Unicode code points orders a supplementary-plane string differently from one in U+E000 through U+FFFF, so two otherwise-conforming verifiers could split on whether identical bytes are canonical. The vocabulary arrays previously said only "sorted ascending" with the order undefined, which is the same crack one clause over from the covering payload surface it was reported against. Pin the labels/caught sort to UTF-16 code-unit order explicitly, and require BMP-only strings on every signed canonical surface (covering payload member names and both vocabulary arrays), rejected as malformed rather than left as producer hygiene: within the BMP the two orders coincide, so the verifier split is unconstructible for conforming implementations. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
Reading reserved members out of a producer-defined signed payload is the JWT registered-claims pattern (RFC 7519), applied inside an attestation token by EAT (RFC 9711), with OCI annotation prefix reservation and the RFC 6839 structured-syntax suffix as the parsing license. Cite that lineage as an informative note at the reserved-member definition, mark the citations as locating the pattern rather than importing any cited standard's rules, and state the two deliberate departures: fail-closed handling of colliding or unrecognized reserved members where JWT ignores unknown claims, and the normative verify-then-read discipline that closes the parse-before-verify deployment mistake. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
Folding a value that was unpredictable before its round, such as a drand round output or an RFC 9334 Section 10.3 epoch identifier, gives the arming record a proven earliest-possible signing time at no new wire member: a signature over a beacon value cannot predate the beacon round. State the recommendation with its three necessary qualifications, since a naive adoption would fake, vacate, or break the floor: the public value augments and never replaces the substrate-unique run-start component; the floor bounds recency only where consumer policy couples the folded round to its freshness window, because the producer selects the round; and the value is fetched at arming time rather than cached, since a stale round folded as current defeats that same coupling. The issued-at timestamp stays the asserted ceiling, deliberately not upgraded to a two-sided proof, and a beacon inside the producer's own trust domain yields no floor against that producer. Public rounds additionally make reuse observations from independent consumers comparable on a shared time axis. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
No self-contained attestation can prove the absence of sibling runs; that is the split-view problem, and fork consistency is its known ceiling. The same construction that makes observation loss gap-evident inside a run, the sealed record's drop count and the checkpoint sequence rule, lifts to runs: an arming payload may carry a run sequence number, the predecessor run's binding digest, and a required chain scope, making a skipped run a gap, a rerun a fork, and a chain reset a duplicated genesis of the same grade as a shared predecessor. The members are syntax-checked only inside one attestation; every rule over them is consumer policy across a published set. The text states the honest limits: ordering under the substrate key only, with commit-before-outcome available solely through the run-entropy floor or an external registration receipt; a numeric gap is unexplained absence and never fraud evidence, with the consumer's remedy being a policy that demands a contiguous fork-free chain; and an unscoped or globally scoped counter is called out as vacuous or as a run-volume leak respectively. TUF's snapshot role signs a population; SCITT registration (RFC 9943) with COSE receipts (RFC 9942) is the consumer-side completion that upgrades gap-evidence to auditable non-omission, cited as the deliberately-external machinery. The run-population non-claim in the strength-orderings section is amended to point at the members without weakening the non-claim itself. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
The shared-index allowance conditioned the gate verb covers on whether a committed payload evidences each referenced attack, an unevaluable predicate, and the selector-absent phrasing implied selectors do covering work. Both leak unevaluable conditions into a section whose headline is coverage as pure byte functions: an aggressive but defensible reader could build an evidencing heuristic into a verifier and split from one that does not. Restate the rule as a producer MUST outside every gate, state explicitly that no validity requirement, recompute input, or tier evaluation reads it and that selector presence changes no gate outcome, and record the forward rule in the changelog as versioning discipline: a member is born exactly when a normative reader consumes it, so if the obligation ever becomes checkable, attribution strength becomes a required member at that version and never retroactively. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
The informative verifier ordering read as a flat list, and two careful readers of the same text counted its steps differently. Restate it as two explicit stages: four numbered byte-pure validity steps that are consumption preconditions, then the trust-relative stage of signature, tier, orderings, and consumer policy. The verification prose also assumed a consumer-pinned expected corpus and substrate without ever stating the obligation. State it under a new consumer-policy-obligations subsection: the consumer pins both digests out of band, compares at consumption, and does not admit on mismatch, with the comparison deliberately kept out of the validity gates because expectations differ per consumer while validity holds identically for all. Recommend that verification surfaces expose a single conjoined admission result so a result-only consumer cannot misread a valid-but-wrong-context attestation as admissible. Extend the 0.6 changelog to cover the full final 0.6 shape. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
The chain-member paragraph defined the syntax rules but left the violation consequence implicit, and said nothing about a scope or predecessor member appearing without the sequence number. State both: any syntax violation, including a chain member present without the sequence, is handled as any reserved-member violation and the record covers nothing. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
The predicate prose used appositive em-dashes and a dense contrastive cadence that set it apart from every other predicate spec in this repository, none of which use em-dashes. Replace the em-dashes with the parentheses, commas, colons, and sentence breaks the siblings use; restate a handful of contrastive closers as plain declaratives; and vary the changelog entry headers so they are not near-duplicates. Editorial only. No normative requirement changes: every MUST/MUST NOT, vocabulary rule, digest definition, gate step, and reserved-member name is unchanged. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
|
Thank you for this, it's the kind of review that genuinely makes the spec better :) I've taken all six points, five into committed text and one as an acceptance with a question attached. As before, I'm replying after the push, so everything below is committed text and not intent, landed as a commit series ending at the branch head. On reserved members, I've turned your precedent framing into the spec's own defense. I added an informative note at the reserved-member definition citing RFC 7519 registered claims, OCI prefix reservation, and the structured-syntax suffix license from RFC 6839, and I mark it as locating the pattern rather than importing any cited standard's rules. One step further along your own argument, EAT (RFC 9711) is the domain-tightest instance, registered claims inside an attestation token, in the RATS family this predicate already leans on. The in-toto Statement itself reads a reserved You're right that the canonicality crack is real, I've adopted the fix, and it reaches further than the surface you named. Folding your fix, I swept the sortedness clauses and found the identical split live between my own rails one clause over. The labels and caught arrays said "sorted ascending" with the order undefined; my TypeScript rail and the Go reference compared UTF-16 code units, while my standalone Python verifier compared code points. I proved it executably: set a supplementary-plane string against one in U+E000 through U+FFFF and they order oppositely, so the same bytes read attestation-valid on two rails and attestation-invalid on a third. My committed fix does both halves. I pinned the sort to UTF-16 code-unit order explicitly, and I promoted BMP-only from producer hygiene to a verifier rejection obligation on every signed canonical surface, meaning member names and both vocabulary arrays. I upgraded it to a reject deliberately, because hygiene still lets an adversarial producer mint payloads that split a code-point verifier from a UTF-16 verifier, whereas the reject makes the divergence unconstructible for every conforming implementation, which is the property your fix was after. That leaves one honest consequence for the vectors you will check against. Once I enforce BMP-only I cannot construct an accept-side probe for the sort order, because inside the BMP the two orders coincide, which is the whole point. So I carry only the reject side, a supplementary-plane member name and a supplementary-plane vocabulary entry, and I pin the comparator regression in per-rail unit tests, the one layer where it stays expressible. On runEntropy, I adopted your idea with credit, as a SHOULD, and I added three sharpenings to keep the floor honest. The folded value must be unpredictable before its round, or the floor is fake. It has to augment and never replace the substrate-unique run-start component, or the anti-splice property collapses into whatever the beacon publishes. And I require it fetched at arming time with the round reference recoverable from the arming payload, because the producer selects the round, so an uncoupled or cached round proves age rather than recency and the floor bounds freshness only where consumer policy couples the round to its window. I say plainly in the text what this guarantees and what it does not: a proven floor, an asserted ceiling, deliberately not a two-sided proof, and no floor at all against a producer whose beacon sits inside its own trust domain. Your comparability point survives all three qualifications and I kept it in, since public rounds give independent consumers a shared time axis for reuse observations. On run population, you're right that the construction lifts, and I carried it into the committed text as the three optional arming members. Two are yours, the run sequence number and the previous-run binding. The third your sketch did not have: I made chain scope a required member whenever the sequence is present. Without a declared scope every rule over the pair goes vacuous, because a producer that chooses its population freely makes any gap uninterpretable, and the obvious default of one global per-key counter leaks the producer's total run volume across customers, so I recommend a minimum of substrate key by subject. I close two more edges. I treat two genesis records under one key and scope as equivocation of the same grade as a shared predecessor, because otherwise the cherry-picker's cheapest move is a chain reset where every winner is sequence one. And I keep the claim language ordering-only, because nothing on the wire anchors when an arming record was signed relative to the run's outcome, so I claim commit-before-outcome only in combination with the beacon floor above or an external registration receipt. I state a numeric gap as unexplained absence and never as fraud evidence, since crashed and private runs produce gaps innocently; what the pair buys is demand-disclosure, where a consumer policy may require a contiguous, fork-free chain over the runs offered to it, fork consistency being the ceiling, as you said. I pulled RFC 9943 and RFC 9942 from the editor and read them before citing; I cite them as the registration completion, with TUF's snapshot role as precedent for the concept of signing a population, and not for this mechanism. If the recommended-minimum scope granularity strikes you as wrong in either direction, that is the one member of the trio I would most value a second reading on. On containmentObserved, your null search stands and I've written the tripwire down. I put it in the changelog as versioning discipline, not in the member paragraph, so it binds future versions instead of decorating this one: a member is born exactly when a normative reader consumes it. Your near-miss was nearer than you called it. The shared-selector allowance as I had committed it wore the gate verb, "covers each referencing row only where its committed payload evidences each referenced attack," an unevaluable condition I had left sitting inside the section whose headline is coverage as pure byte functions. Your from-spec checker would have had to pick a reading there. So I reclassified it. I now say a producer MUST NOT reference a record from a row its payload does not evidence, stated as a producer obligation outside every gate, and I say explicitly that no validity requirement, recompute input, or tier evaluation reads it and that selector presence changes no gate outcome. Thank you for offering to build the checker, and yes, I'd love that! Here's what you can hold me to: the suite is deterministic, 34 accept and 91 reject vectors as of this revision's additions, every vector a bare Statement with test keys you can derive from a published seed formula, and per-vector traceability from condition to spec line. So that the experiment stays honest, I declare in the manifest that the verdict, and each accept's result token, is the normative comparison surface and that the per-vector condition codes are informative, since the code set is my implementation's vocabulary and pinning it would rig the uniqueness question. And I numbered the validity steps in the spec, because your "four" against the text's five bullets was two careful readers counting differently, which is itself a legibility datum. I read your four as the byte-pure pipeline partition, statement well-formedness, coverage validity, the recompute, and digest integrity, with the tier excluded as the one trust-relative stage; confirm that is the scope you intend to implement. My one ask is that you have your checker report a free-form reason per reject alongside the boolean, mapped to my codes only informatively. Verdict-only scoring would let a wrong-reason reject score as parity on most of the corpus, and I care about the divergences where two implementations reject the same vector for different reasons, which are the actual product of the exercise. And a gist would be the wrong home for it: land the checker in or alongside the suite under your own authorship when it exists, since a two-independent-authors record is worth more durably than a comment thread. It's public now, Apache-2.0, at github.com/astrogilda/aee-conformance, with a MANIFEST enumerating each vector's expected verdict and result, the Go reference verifier, a standalone Python rail, and a harness that replays the whole corpus. That is the place to land it; the vectors and the spec text above are committed there and on the branch head. On the verification story you reviewed, one more committed change. I now state the anchors a consumer pins, the expected corpus and substrate digests, as explicit consumer-policy obligations with a single conjoined admission result, and I keep them deliberately outside the byte-pure gates: validity holds identically for every consumer, and which corpus you meant to assess against does not. Thanks again for the care in this round. The two things I'd most value your read on are the validity-stage scope you intend to implement and whether the chain-scope granularity lands right, but honestly the whole pass has been a pleasure to work through. Edited 2026-08-31: prose only, for readability. No figure, vector count, citation or claim changed. |
|
Scope confirmed as you numbered it: the checker implements the four byte-pure stages, statement well-formedness, coverage validity, the recompute, and digest integrity, from the spec text at the branch head, with the signature tier excluded from the parity claim. One boundary note so the corpus and the claim line up cleanly. Where a reject vector is reachable only through the tier, I will still run the DSSE verification against the seed-derived test keys, because a test corpus pins its anchors and pinned anchors make the tier deterministic; the trust-relativity lives in anchor choice, and for the vectors the manifest has already chosen. So the parity claim is stages one through four from spec text alone, with the tier exercised under the corpus's own keys rather than reimplemented as policy. Reason-per-reject is the right experimental design and I would have asked for it if you had not: verdict-only scoring lets a wrong-reason reject pass as agreement, and the divergences are the product. On where it lands, I am taking the alongside arm of your offer: an own repository under my authorship that the suite references, since a second implementation is most legible as a second implementation when it has its own home, history, and CI, and a one-line link from your README is all the coupling the record needs. If you would rather have it in-tree, or would rather the checker not touch DSSE at all, say so and I will follow the suite's lead on both. On chain scope, the floor is right, and the member is the one of the three that needs one more property to do its job. Key by subject is exactly the consumer's question, whether this image was rerun, and the per-key default you rejected does leak total run volume across customers the way you said. The gap is that a recommended minimum cannot be gated, and your own reset rule tells you where the pressure goes next. Once a chain reset is equivocation, the cherry-picker's cheapest move migrates to scope narrowing: declare scope as key by subject by corpus by posture, and every published run is legitimately genesis one of its own chain, no gap, no fork, nothing to detect. Demand-disclosure binds only if a consumer can check that the declared scope is no finer than the scope its policy demands, and that comparison needs the member to be machine-comparable rather than free-form: a closed set of dimension tokens with a pinned order, compared by subset, so that "no finer than substrate key and subject" is as mechanical as the rest of the validity checks, with an unrecognized token failing closed the way every other closed vocabulary in this spec now does. One reading is worth pinning in the member's definition while you are there, because the two halves of the design read different things. The policy comparison reads the declared dimension set; the equivocation rule has to key on the evaluated tuple, this key and this subject's values, not the token set. Under that reading genesis-per-subject is the normal case rather than a reset, and under the token-set reading every second subject's genesis would collide as equivocation, so the sentence is load-bearing rather than pedantic. With both in place the narrowing move becomes as visible as a gap or a fork, and your recommended minimum can stay a recommendation, because the policy hook is what does the binding. Everything else in the round reads settled to me, the lineage note locating rather than importing and the changelog tripwire binding future versions where it belongs. I will start from the numbered spec text and the manifest, and the vetting second-reviewer commitment stands. |
|
One correction to my last message, and a note on the fix I promised in it. I wrote that three vectors carry an arming posture differing from their pinned digest: The row-level scope fix has landed. Adversarial review of it caught something worth passing on, because it is the same shape this thread keeps turning up: the arming term had three consumers, not two. It rode inside one boolean and so reached the row check, the statement-level existential and the universal carried-record sweep at once. Hoisting it out to make it row-dependent gave it back to two of them and silently dropped it at the sweep, whose refusal message went on telling producers that the posture had been compared against every carried arming record. Restored, and the message is true again. Neither the corpus nor our 250 could see any of it. The changed line is evaluated 70 times and none of those 70 can tell the per-row set from the union, so the fix is untested here rather than confirmed. The vector that would discriminate is a clean row resolving a valid arming record and a valid seal, while a different row resolves one at a divergent posture. |
The coverage-partition rule (spec L381-383) makes the three sets a disjoint partition of the manifest's classes, so membership runs both ways, but only bad-819 forced the assessedClasses side. Add bad-731-outofscope-unknown-class and bad-732-routedelsewhere-unknown-class: an unknown class key in each reason map, result left alone, rejected coverage-incomplete. Both reference rails already enforced it; the vectors lock the rule and mutation-prove the rails. suiteRevision 3, 140 vectors; registry decision 14 extended; docs record the registry as a post-run reconciliation surface. Surfaced by the independent from-spec checker (in-toto/attestation#570 round-8).
The vendored copy of the predicate specification was pinned at 23bee586, which predates the paragraph requiring every carried record of a covering kind to satisfy that kind's constraints whether or not a row resolves an index to it. suiteRevision 23 shipped sixteen reject vectors against that requirement, so the suite refused sixteen statements on a sentence a reader of the bundled copy could not find. An implementer who pins this bundle and implements from it alone, which is the reason the copy is vendored at all, would read those refusals as the corpus overreaching its own text. Re-vendor to c0c4da67defdf0f186f162e7ecb3f9527b6a94f8, the current head of in-toto/attestation#570, via scripts/vendor-spec.py. The specification gains the missing paragraph and its upstream changelog entry and nothing else; the line citations and anchors are remapped mechanically and re-pinned. The corpus does not move: 248 vectors, the same corpusDigest, and both reference rails still pass 248 of 248. Record the move as suiteRevision 24 and carry the number into the documents that restate it. The independent checker has posted no run at this revision, so it joins the list of revisions the report publishes no score for.
|
Taking the arming-scope question first, since you left it with me: yes, it earns a vector, and your own last message is the argument for it. You shipped the row-level scope fix and then said the changed line is evaluated seventy times and none of those seventy can tell the per-row set from the union, so the fix is untested here rather than confirmed. That is the case for adding the shape. A rail whose only discriminating statement is So the corpus will carry a statement with every row caught and two rows resolving different subsets of two arming records at divergent postures, which is the shape you described. I will also add the On the census retraction, and I mean this as more than politeness. Testing whether the decoded payload was an object without decoding it, and getting zero everywhere, is the third instance of one shape in this thread: a zero that is structural reading exactly like a zero that is empirical. Your The correction from three vectors to two is right, and Your three-consumers finding is worth passing back to you. The arming term reaching the row check, the statement-level existential and the universal carried-record sweep inside one boolean, then getting hoisted and silently dropped at the sweep while the refusal message went on advertising a comparison it no longer performed, is the most useful thing either of us has found this round. A refusal message that overstates what was checked is worse than no message, because it is the thing a producer reads to decide they are done. I would rather the specification said explicitly what a refusal is allowed to claim than leave it to each implementation to keep its message honest under refactoring. Edited 2026-08-31: prose and paragraph spacing only. No figure, vector name or claim changed. |
The type URI appeared in five places as a bare fact, so a reader who followed it got a 404 with nothing in the repository explaining why, and the reasonable conclusion from that is that the specification does not exist. The in-toto attestation catalog redirects the URIs of vetted predicates whose specification is merged. This predicate is in review as in-toto/attestation#570, so its URI is not yet served, which is the ordinary condition of a predicate at that stage. Each of the four prose sites now states that, says the URI identifies the predicate type rather than being fetched during verification, and points at the vendored copy of the specification that this repository carries. The fifth site is the vendored specification itself. It stays byte-verbatim: 121 spec:NNN citations index into it by line number and its SHA-256 is pinned in VENDOR-PIN.json and MANIFEST.json, so the note for it lives in spec/README.md, which introduces the copy and covers the schema $id alongside it.
|
A correction I owe you from before you started building, and then the vector. When I handed the suite over on 23 July I told you I declare in the manifest that the verdict, and each accept's result token, is the normative comparison surface, and that the per-vector condition codes are informative. For a month the manifest carried no such declaration. Its keys were suite, predicateType, specPath, specDigest, tracksUpstream, specUpstreamCommit, counts, corpusDigest and vectors, and On the arming scope I committed to a statement with every row caught and two rows resolving different subsets of two arming records at divergent postures. I built both halves with the corpus's own generator and measured them on five compiled rails. The every-row-caught half is inert. It produces byte-identical output on every rail I can construct, including the one with the conjunct deleted by the corpus's own mutation tool, and bad-902 moves under that mutant, so the read path is proven. The mechanism is that The other half survives and does more than I promised, once it is paired with a second vector I had to invent. Both are built and measured here and neither is on main yet, so do not go looking for them in the corpus: they will land as bad-1018, carrying a divergent arming record no row resolves, and bad-1019, carrying one resolved only by a caught row. Read together they name the reading: both report One reason the corpus could not see any of this is mine. bad-703 carries the same divergent-arming shape, and its expected set names all three of arming-covers-nothing, sealed-covers-nothing and clean-row-uncovered, so it is satisfied by any of them and stops measuring the question instead of starting to. The two new rejects carry single-valued code sets, and the On your three-consumers finding, here is what a specification sentence has to constrain, since I prefer writing it to admiring it. A refusal message is itself a claim about which comparisons ran, so bind it to the evaluated set: an implementation can name a condition in a refusal only where that condition was evaluated on that statement, and it cannot name a comparison whose operand set was empty. Your hoist satisfies the first clause and breaks the second, which is the part that interests me: the sweep still ran and still refused, and what changed is that the arming term arrived carrying nothing while the message went on describing a comparison against every carried arming record. Written that way the constraint is checkable from outside a rail, since the reported condition set becomes a function of what was evaluated rather than of what the author put in the string. The case I am least sure of is a comparison evaluated over a set the implementation cannot enumerate at message time. Did you hit that in the checker? Edited 2026-08-20: this named two vectors, a reading entry and a ledger cell in the present tense when all four are built here and not yet on main. Corrected above so nothing sends you looking for something that is not published. Edited 2026-08-31: prose only, for readability. No vector name, figure or claim changed. |
|
Two corrections matter if you are reading from this pull request. Both concern the gap between the head here and my vendored copy. I have edited the comments they came from, so the wrong versions no longer stand above the fixes. On 11 August I wrote that the reporting paragraph was in the branch, and it is not. The head here is 95470f3 and it carries none of that wording; the coverage sentence still reads "the arming record's" at line 1349, so the definite-singular fix I described in the same message is also only in my copy. Everything I said about what those edits do stands, and none of it has landed upstream yet. The And from this morning's message, bad-1018 and bad-1019 are built here and not published, along with the reading entry and the ledger cell that go with them, so a grep for any of the four in the corpus will correctly find nothing. Edited 2026-08-31: prose only, for readability. No commit, digest, line number or vector name changed. |
|
Yes, three times in one message, and none of them is the case you are least sure of. Applying your The site is the carried-record sweep,
It names comparisons that did not run. The conjunction is computed in It names a comparison whose operand set is empty a third of the time. It names a set wider than the one compared, and wider than your spec. This is the one I would The honest bound on all three, because it changes what you should do with them. These are One consequence I did not enjoy finding: my comment claiming the restoration "keeps this site's On what you actually asked. We did not hit it. At message time Where your uncertain case does live in ours is not a set but a commitment. Which suggests emptiness is the special case rather than the rule. The general form your two clauses Last, the measurement bound that explains why neither of us saw any of it. No vector detects any of All three are fixed, with the message assembly moved into a pure function so a rule the file states One result from that work belongs to you rather than to me, because it is your inert half arrived at One narrowing on your own correction, since it changes what a third party can check. You wrote that One retrieval caveat I hit while re-running, and it is the same trap one commit further out. Suite The fix and its tests are at Rul1an/aee-checker#16, with the bite results in the description. AI-assisted; I ran the reads, the runs and the mutations, and am responsible for the claims. |
|
@astrogilda asked me to read the substrate half with a custody eye. Against 0dbe10b. Section references are to "Three Jobs, Not One" (Zenodo 10.5281/zenodo.21935891 v6), which sets out the threat model this predicate is written against. Two points:
Both emit "attested". Even with a pinned key, the attestation itself says nothing about where that key sits. A consumer that provisioned the deployment already knows; one that did not — the third-party reader Definition 6 is written for — cannot tell whether the evidence came from a monitor your model supports or one it defeats. Not asking for a new field — you already carry two closed axes and a third costs real complexity. Whether this belongs on the wire, in consumer key policy, or simply as a stated limitation is your call. |
Revise the AI Agent Action predicate based on detailed review feedback: - Add checkpoint and chain_break as action.type values with sub-schemas, addressing the tail truncation detection gap - Add parties array with witness/asserter roles for field provenance - Split canonicalization: signing tuple-array (M/L-tagged) for chain integrity, RFC 8785 JCS for content digests (float-safe for MCP payloads) - Document genesis convention (previousHash: "genesis") and chain_break requirement for crash recovery - Add I-JSON safe integer bound on all integer fields (RFC 7493) - Add 128-level depth bound on extensions with counting rule per in-toto#570 - Replace placeholder digests with fully recomputable worked example (every digest has a shown preimage, verifiable against conformance vectors) - Add Security Considerations section - Move listing to community contributions per ITE-63 process
Revise the AI Agent Action predicate based on detailed review feedback: - Add checkpoint and chain_break as action.type values with sub-schemas, addressing the tail truncation detection gap - Add parties array with witness/asserter roles for field provenance - Split canonicalization: signing tuple-array (M/L-tagged) for chain integrity, RFC 8785 JCS for content digests (float-safe for MCP payloads) - Document genesis convention (previousHash: "genesis") and chain_break requirement for crash recovery - Add I-JSON safe integer bound on all integer fields (RFC 7493) - Add 128-level depth bound on extensions with counting rule per in-toto#570 - Replace placeholder digests with fully recomputable worked example (every digest has a shown preimage, verifiable against conformance vectors) - Add Security Considerations section - Move listing to community contributions per ITE-63 process
Revise the AI Agent Action predicate based on detailed review feedback: - Add checkpoint and chain_break as action.type values with sub-schemas, addressing the tail truncation detection gap - Add parties array with witness/asserter roles for field provenance - Split canonicalization: signing tuple-array (M/L-tagged) for chain integrity, RFC 8785 JCS for content digests (float-safe for MCP payloads) - Document genesis convention (previousHash: "genesis") and chain_break requirement for crash recovery - Add I-JSON safe integer bound on all integer fields (RFC 7493) - Add 128-level depth bound on extensions with counting rule per in-toto#570 - Replace placeholder digests with fully recomputable worked example (every digest has a shown preimage, verifiable against conformance vectors) - Add Security Considerations section - Move listing to community contributions per ITE-63 process
Three corrections to normative text, all raised in review by two reviewers. The first gives the sealed kind constraint the equality its arming sibling already carries. The arming bullet requires aeePostureDigest equal to the pinned networkPosture digest; the sealed bullet required the member and stated no equality at all, leaving the seal's run-end posture compared to nothing on a check that reads no row. The equality existed only inside a row-scoped sentence, and that sentence says outright that which records supply its set on a check reading no row is not settled by it. Arming was already settled on that path, because its bullet pins every arming record's digest and kind constraints are evaluated over every carried record binding to the run whether or not a row resolves an index to it, with a violating record covering nothing. The seal now settles the same way, by the same mechanism and in the same words. The second corrects a count that the same release made stale. The sentence separating this predicate's axes from producer territory named two of them, basis and method. Attribution is now REQUIRED on every row over a closed vocabulary, and the text above it already says attribution is the axis that acquires a normative reader at this version, so the criterion the sentence exists to state selects three. The criterion is unchanged; only the count and the naming follow it. The third adds a residual the paper stated and this document did not. A monitor supervising syscalls from the host kernel and a monitor reading guest state from outside a virtual machine both satisfy the requirement that the observation key not be accessible to the subject artifact, both sit at basis substrate because that axis names a class of vantage rather than the identity of the substrate holding it, and both derive attested, a tier that is a function of a signature verifying against a policy-named key and reads nothing about where the key is held. A consumer that did not provision the deployment cannot tell the two apart. The residual says so and points at the refinement path already carried under the evidence tier, where consumer policy MAY subdivide attested into stricter refinements. No member, axis or vocabulary is added, which is the point: the distinction is bought with key policy. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
|
@Rul1an, @zlhk100: both of you are right, and in one case righter than you claimed. Rul1anYou are right on all three. The third one is aimed at my text, so I checked it against the tree. "carried arming" is not in the specification. Zero hits across all four files this PR changes, on three search shapes: the literal phrase, a loose The narrowing you concede is the one my text already makes. The equality in my sentence is scoped at SPEC:1364-65 to "the For arming records the set question is already moot on any valid statement. SPEC:1327-28 requires every arming payload to carry an aeePostureDigest equal to the pinned It is the seal where this is open, and that is a real hole. Compare the two kind constraints: the arming bullet carries its equality clause inline at SPEC:1327-28, and the On your three reporting defects: I have no argument with any of them, and the short-circuit one is the same complaint I made about one name covering three kinds, one level down and in my direction. The 22-of-66 empty-operand measurement is the kind of thing I should have asked for and did not. zlhk100Right on all four legs, and the fourth is the one that costs us something. §3.1 does grant arbitrary system calls beneath any guard and does name the kernel interfaces beneath the host process among attackable surfaces. The trust-domain definition then licenses your inference directly: a process, privilege or address-space boundary between a component and a record "does not alter the result", and components in one trust domain fail together. §4.5 puts it more strongly than you did. "A kernel privilege escalation collapses the observer and the observed into one trust domain", and "Once a workload reaches Ring 0, it also reaches the observer". Your quote drops the flattering half of that sentence. On the wire your reading is exact. There are three closed axes now, where you counted two: Appendix E of the paper already states your finding as a general rule: "A reader who cannot tell which method produced a record must treat the record's independence claim as unproven." §4.4 levels the same complaint at a commercial product and calls it "the check a buyer should ask for". So you have not found a gap in the argument; you have found the argument applied to my own wire format, which is worse for me and better for the document. The asymmetry worth naming: the spec states the non-identification as a design intention at SPEC:998-1002, and its residuals section at SPEC:663-760 lists four things the 0.7 commitments do not close, none of which is vantage-class opacity. The paper says it three times and the spec never says it as a limitation. That is the fix I owe, and I agree with you that it does not need a new field to make. You said the placement is my call. I took the third option you offered, a stated limitation in the residuals list, plus a pointer to the refinement path the spec already carries at SPEC:775-78, where consumer policy MAY subdivide attested into stricter refinements such as requiring a hardware-attested observation key. That names the escape without spending an axis on it. Pushed as a4cb887
None of it needs a new member on any closed vocabulary. The residual says why a carried value would not help: a vantage-class member would be a producer assertion about the producer's own stack, exactly as forgeable as the rest of the payload, and it would state at the predicate level a thing the consumer's key policy already decides. Edited 2026-08-25. The quotation of the Edited 2026-08-31: prose and formatting only. No SPEC line reference, quotation, commit or claim changed; some repeated field names lost their backticks after first mention. |
|
Re-read The seal side is closed. With the equality now inside the Your zero-hit result for "carried arming" is also right: the phrase was ours, not a quotation from this specification. We used it to name the wider quantifier reading while rejecting it, and recorded that the corpus could not distinguish that reading from the narrower row-resolved one. The current checker no longer emits that phrase after the reporting fix; it remains only in historical, digest-pinned records and explanatory tests/comments. Your narrowing of the arming half stands, and we had not claimed otherwise. Thanks for checking this against the tree rather than recollection. |
A refusal names a comparison, and nothing until now required that the set it names be the set the implementation actually ranged over. Three shapes break that, and they break it differently, so the clause states each with its own remedy rather than as one caveat. A comparison whose operand set was empty did not run. The document already argues this twenty lines above, where the attribution requirement exists precisely because a requirement universally quantified over an empty set is vacuously true, so the clause cites that argument rather than repeating it: a verifier must not name such a comparison, and should report the empty set, which is what it found. A set described wider than the one the check ranged over is the shape that costs a producer real work, because it tells them to repair records the check never read; where the evaluated set is a function of the statement, as the set of arming records a row resolves is, the clause asks for that set and its count. And a comparison against a commitment genuinely ran, but has no exhibitable membership at all: a digest mismatch says the recompute differs and says nothing about which element differs, so a refusal describing it as a comparison over a set reports a capability the construction denies it. The obligation is diagnostic and never a validity rule, and it sits in the paragraph that already establishes that hedge so it inherits it. No conformance vector can enforce it. Deleting an emptiness guard leaves a suite green because the enclosing universal is vacuously true on the empty set, which is the same structural zero the clause describes; and this document defines no condition vocabulary, so what a refusal is drawn under is not fixed here either. What is fixed is only what a refusal may claim to have compared. Signed-off-by: Sankalp Gilda <sankalp.gilda@gmail.com>
Revise the AI Agent Action predicate based on detailed review feedback: - Add checkpoint and chain_break as action.type values with sub-schemas, addressing the tail truncation detection gap - Add parties array with witness/asserter roles for field provenance - Split canonicalization: signing tuple-array (M/L-tagged) for chain integrity, RFC 8785 JCS for content digests (float-safe for MCP payloads) - Document genesis convention (previousHash: "genesis") and chain_break requirement for crash recovery - Add I-JSON safe integer bound on all integer fields (RFC 7493) - Add 128-level depth bound on extensions with counting rule per in-toto#570 - Replace placeholder digests with fully recomputable worked example (every digest has a shown preimage, verifiable against conformance vectors) - Add Security Considerations section - Move listing to community contributions per ITE-63 process Signed-off-by: Elan Ansrinivasan <5340827+elang2@users.noreply.github.com>
Revise the AI Agent Action predicate based on detailed review feedback: - Add checkpoint and chain_break as action.type values with sub-schemas, addressing the tail truncation detection gap - Add parties array with witness/asserter roles for field provenance - Split canonicalization: signing tuple-array (M/L-tagged) for chain integrity, RFC 8785 JCS for content digests (float-safe for MCP payloads) - Document genesis convention (previousHash: "genesis") and chain_break requirement for crash recovery - Add I-JSON safe integer bound on all integer fields (RFC 7493) - Add 128-level depth bound on extensions with counting rule per in-toto#570 - Replace placeholder digests with fully recomputable worked example (every digest has a shown preimage, verifiable against conformance vectors) - Add Security Considerations section - Move listing to community contributions per ITE-63 process Signed-off-by: Elan Ansrinivasan <5340827+elang2@users.noreply.github.com>
VENDOR-PIN.json named in-toto/attestation and left ref empty. The commit it pins is not in that repository: a plain clone of it resolves neither 0dbe10bc nor 639ec56c, while the same clone resolves its own HEAD, and a clone of astrogilda/attestation resolves 0dbe10bc on branch predicate/adversarial-execution-evidence, where it is the direct parent of a4cb887 and its spec file is 147709 bytes hashing 759d2383, byte-exactly the pin's own specDigest. So the digest was right and the address was wrong, and a reproducer following the pin arrived at a repository the bytes are not in. The cause is one field answering two questions. A pull request is REVIEWED in the upstream repository and opened FROM a branch in a fork, and upstreamRepo was being read as both. It stays as the review venue, because gen_manifest.py builds the citation in-toto/attestation#570 out of it and that citation is correct. Where to fetch becomes commitRepo, ref and refKind, and all three are derived from the checkout's remotes rather than typed, for the same reason the commit already was. vendor-spec.py now refuses to write a pin whose commit is not reachable from the ref it names. That refusal fired on its first run against a real tree: the local branch is one unpushed commit ahead of the fork, so vendoring from it would have pinned a commit no reproducer could fetch. --at separates the commit vendored from the ref that contains it, which is the situation here, since the corpus certifies against an ancestor of the branch tip. refKind is the part that says the pin is currently-true rather than permanent. A branch head moves and this one already has. The tag that fixes that is a remote write and is left for the operator, with the commands in TODO.md. Dry-running those commands first is how the annotated-tag peel bug surfaced: rev-parse on an annotated tag returns the tag object, git show dereferences it, so the digest check passed while the pin recorded an id that is not a commit. The two documents that quote the provenance are regenerated from their generators. Both said "upstream commit <sha>" beside a tracksUpstream of in-toto/attestation#570, which reads as an instruction to fetch from there; both now name the fetchable location and the review venue separately. Staging the pin in the forcing gate's test rig is the other half. The gate reads a file the rig did not copy, so every rendering case died on a missing file and reported that as its own verdict: the absent-prior case failed saying the prior record had not stopped the run, when the run had stopped one file earlier. The gate also refuses now instead of raising, because a traceback exits 1, which is this gate's code for a stale document.
|
Thanks — the limitation as written reads correctly, and the vantage-class decline is a reasonable place to land it. One data point you may already know, in case it's useful. SLSA meets both of your objections and lands the other way: builder.id is a producer-supplied string naming the platform rather than the key, and the spec keeps it REQUIRED even if it is implicit from the signer. It also asks that modes with differing security attributes each carry a different id — GitHub-hosted versus self-hosted runners is the worked example — to minimize the risk that a less secure mode compromises a more secure one. Not arguing it's the better call here; builder.id is self-declared and carries the same forgeability you name. Only that neither forgeability nor overlap with key policy settled it in that predicate, which I found interesting given the family. Happy to leave it there. Good luck with the vetting. |
…JSON Adopts the canonicalization and chain-shape replacement prose contributed in review by Sankalp Gilda (@astrogilda), adapted from the text of in-toto#570. - bad-101..bad-104: the chain hash preimage is now the record canonical form: RFC 8785 (JCS) over the complete audit record including its attestation member. Producers MUST write the JSONL line as exactly these bytes; verifiers MUST recompute and reject on byte mismatch, fail-closed. JSON.stringify is banned from deriving the record canonical form. - bad-105: strict I-JSON statement-wide. Duplicate members at any depth reject fail-closed; well-formed-string rules on raw bytes and \u escapes; the 128-level depth bound now covers the whole record. - bad-107: checkpoint linkage wired. Every record type carries predicate.chain.previousHash; checkpoint.previousHash restates the head and MUST equal it; deleting a checkpoint now breaks documented linkage. - Unnumbered findings: previousHash constrained to lowercase 64-hex or the literal genesis; second-genesis detection upgraded SHOULD -> MUST reject; content digest preimages pinned (params member for requests, result member for success responses, error member for error responses); floats permitted in content payloads only; the signing-form field list enumerated in the spec text for all three record types. - Widened the tool_call signing tuple to cover type, errorClass, the content digests, attestorVersion, and configHash (adopting the reviewer's proposed additions), and added the Underlying record shape section pinning record-vs-Statement membership and the three protection layers (signing tuple, chain hash, DSSE envelope). - Hardening from adversarial self-review: RFC 2119 requirements notation; attestation byte encoding pinned (lowercase 64-hex HMAC-SHA256, 128-hex Ed25519); malformed JSON-RPC responses recorded with success false and no response digest rather than suppressed; predicate.chain REQUIRED on tool_call and checkpoint records (null-priorHead chain_break is the single legitimate omission); chain_break prior* members always present, explicit null when unknown; id pinned as opaque, at most 128 bytes, unique per chain; JCS string rules applied universally in the signing form; protocol and upstream.transport declared Statement-only, DSSE-covered annotations with fixed v0.1 vocabularies. - Conformance corpus subsection pinning astrogilda/aee-conformance, with the conformance MUST binding at the pin regenerated against this text. - Security Considerations: planted-break residual risk named with three implementable mitigations; signature-scheme trust model documented (Ed25519 for attestor accountability, HMAC-SHA256 only inside a single trust boundary). - Extensions retention pinned: post-signing stripping is forbidden (it corrupts the chain preimage); a never-inlined workflow emits extensionsDigest with extensions absent at emission time. Scope semantics clarified: scopes describe witnessing, not signing; structural members are emitter-witnessed by construction; unscoped fields remain unknown provenance, fail-closed. - Markdownlint conformance against the repository config: signing-form field lists moved to fenced blocks, fence languages and list-marker spacing normalized. Checkpoint and chain_break Statement schemas now show the metadata block their worked examples carry, and decisionContextDigest is pinned to lowercase 64-hex with a deployment-defined preimage at v0.1. - Worked example recomputed under the record canonical form and widened tuple, with subject digests fixed to the genesis chain hash on every statement in the chain. Co-authored-by: Sankalp Gilda <23521054+astrogilda@users.noreply.github.com> Signed-off-by: Elankumaran Srinivasan <5340827+elang2@users.noreply.github.com>
|
Rul1an, four things have landed since your last read: a tag, a repository name in the vendor pin, a digest pin, and the refusal-set clause. I announced none of them. A silent push and no push look identical from outside. The checks are written out below so you can run them yourself. Start with the suite commit that sat on no ref. You suggested a ref, if I wanted the pinned revision reachable the way a reproducer reaches things. There is one. It went up the same evening: an annotated tag, cited/5019931. Its message records why. History was rewritten after the revision was published. The commit is an ancestor of nothing. A plain clone still carries it, because clone fetches tags. I checked that the way a reproducer would, and not the way my CI does. A plain On the vendored specification's dangling upstream commit. Your finding was that 0dbe10b resolves in neither in-toto/attestation nor the fork. The real defect was one level up. The vendor pin named a commit. It never named the repository holding it, so a reader had to guess, and the natural guesses fail. be67a74 adds A plain clone of that repository reaches the commit. One retrieval caveat worth passing on, because it gave me a wrong answer first. The GitHub API resolves that same SHA under in-toto/attestation as well, since forks share an object store. It answers 200, not 404. So an API lookup cannot settle the question you actually asked. Only a clone can, and the two disagree in the direction that reassures you. For the agent-action corpus the fix was to stop pinning a commit at all, since its upstream commit was 639ec56. That one is genuinely orphaned. The fork branch was rewritten. A plain clone of elang2/attestation exits 128 on it. That clone's own HEAD resolves. 7aed3cc moves the manifest onto Finally, the refusal-set clause, and this is the part I most want you to check. Commit b1513f6 states the general clause in the specification text. Where a refusal names a comparison, the set it names must be the set the implementation evaluated. All three of your shapes are there. Each has its own remedy, and none is folded into a single caveat. An empty operand set means the comparison did not run. A verifier must not name it. It should report the empty set, which is what it found. A set wider than the one the check ranged over tells a producer to repair records the check never read. A verifier must not name a wider set. Where the evaluated set is a function of the statement, it should name that set and its count. The arming records a row resolves are exactly such a set. A comparison against a commitment genuinely ran, and may be named. It must not be described as a comparison over a set whose membership it cannot exhibit. A digest mismatch says the recompute differs, and says nothing about which element differs. The general form is yours. You wrote that emptiness is the special case, and that the set a refusal describes must be the set the implementation evaluated. I took your sentence over my narrower one, and the commit message says so. The limit is worth stating too. No vector can enforce any of it. Your own bite test found the reason. Deleting an emptiness guard leaves a suite green, because the enclosing universal is vacuously true on the empty set. The obligation is diagnostic and never a validity rule, and it sits in the paragraph that already carries that hedge. One thing I have not done, so you do not go looking for it: the vendored copy in the suite is still pinned at 0dbe10b. It carries neither a4cb887's seal-posture equality nor this clause. The pin is deliberate and its digest matches. It moves when the suite re-vendors, which is a separate change from this one. But if you re-read the vendored text expecting the new paragraph, you will not find it there; it is at the branch head. Sankalp. |
|
@zlhk100, thank you, and the SLSA point is better than a data point. It counters the reasoning I gave, so I want to answer it directly. My residual declines a vantage-class member on two grounds. The value would be a producer assertion about the producer's own stack, as forgeable as the rest of the payload. And the discriminating fact already lives in consumer key policy. You have shown me a predicate in the same family where both hold of The practical difference is in what the field is for. That field is a partition key rather than a graded claim. A consumer's policy matches on it, and SLSA keeps it mandatory even when it is implicit from the signer. That is what makes it usable: the field is there to be keyed on, not to be believed. The guidance you quote about modes with differing security attributes is the same thought, since a partition is useless if two modes share a label. My axes are graded. They feed the evidence tier and the result recompute. A self-declared value in that position would be forgeable in a way that changes a verdict rather than a lookup. So the correction is narrower than a reversal. I still decline a graded vantage axis. But what I actually wrote argued against any carried value at all, and that is wider than my grounds reach. An ungraded partition identifier is a different proposal, it is the one SLSA validates, and I have not evaluated it. So I will narrow the residual's wording to the graded axis it actually rules out. That is a text change in the residuals section, and it adds no member to any vocabulary. None of that changes the limitation itself, which stands either way. A third-party reader still cannot tell a kernel-vantage monitor from a monitor outside the workload. An identifier the producer chooses would not by itself let them. It would no more do so than builder.id lets a SLSA consumer judge a runner they know nothing about. What it would give is something for a policy to key on, which is a smaller and more honest thing to claim. Thank you for reading the substrate half with a custody eye. The comparison to the SLSA family is useful context, and I am glad you raised it. |
|
The Adversarial Execution Evidence predicate is a strong addition to the attestation family: running an untrusted artifact against a known corpus produces evidence that attests observed behavior, not trust in the artifact. The recomputable-evidence question it raises:
Attesting the boundary between allowed and observed is precisely the evidence layer I build: what an agent is authorized to do vs what it actually did, recomputably. I work on that at https://agentkey.us and open to comparing predicate shapes. |
|
Sankalp, I reran these as clone checks rather than API lookups.
The refusal-set clause also found a defect in our checker rather than only confirming the earlier fix. Two On the remaining wording question: I found no occurrence of “short-circuit” or “conjunct” in the clause. Our earlier I did not run the current corpus against this checker or derive a number. The checker targets predicate v0.6; the corpus is v0.7 with no alias or dual-accept window, so the type refusal would make such a figure meaningless. |
Adversarial Execution Evidence predicate
This adds a predicate for the evidence produced by deliberately executing an untrusted artifact against a known adversarial corpus inside a containment substrate: agent tools, MCP servers, plugins, build steps. The producer runs the artifact, injects the corpus, and signs what the substrate intercepted.
The predicate has been through fifteen rounds of review here, and @Rul1an has written a checker from the text alone and runs its conformance corpus, so the field shape is settled enough to read as written. They are the same reader in every reference below. I still want review of the fields before the vetting meeting. To answer the new-predicate guideline questions up front:
The use case is an admission controller gating a third-party MCP or agent image on evidence it was detonated against a named corpus under an enforcing catch policy, and an auditor re-verifying a specific interception offline without trusting the producer's infra. Existing predicates do not cover it: runtime-trace carries raw monitor activity with no corpus binding, no coverage denominator, and no per-event signature; SCAI carries attribute assertions, not an adversarial corpus with recomputable outcomes; VSA and SVR carry policy verdicts computed downstream of evidence like this. This is the evidence layer those consume, so verdict semantics stay out of scope here. The policy question it answers is which attacks this exact image faced, what the substrate did about each, and under what network posture. A consumer recomputes all of that from the attestation plus the producer's published taxonomy, with no call back to the producer.
The shape carries an explicit
doesNotAssertnegative scope, and five further design properties motivate it. I'm happy to defend or change any of them in review:corpus.digest.batchRoot.aeemember prefix is reserved for future versions of this predicate and everything else in an observation payload stays producer territory, but no member of that territory is read by a conforming verifier: a producer-defined member MUST NOT affect structural validity, MUST NOT affectresult, and MUST NOT affect the evidence tier, whether or not its values can be ordered. A producer-defined ordered axis additionally MUST NOT be ranked by a verifier and MUST NOT be composed by weakest input across records or rows, because the two axes this predicate does order,basisandmethod, are ordered only because a normative reader consumes them, and a producer axis acquires no such reader by sharing the envelope. That constraint sits in the paragraph granting the territory, rather than beside each future member, so an implementer meets it before the member exists and not after the first one has shipped. The first version of this rule reached only ordered members. They applied it to their own tooling the day it published and reported two gaps: it missed their unordered producer members, and one of their checkers was gating validity on a producer member's value, which their own design document forbade. Both halves of the wording above come from that report.One thing I'd specifically like the maintainers' read on: whether evidence that carries no verdict belongs as its own predicate or should be framed relative to SCAI.
We build monitors that would emit this downstream of runtime traces. For folks in #557 and #568 working on eBPF and CI monitors (@rung, @stupendoussuperpowers): this is designed to wrap and bound the outputs of a trace, and I'd value knowing whether your tools' trace-policy and scope outputs map cleanly onto
observationEnvironmentand the coverage sets.Vetting status (v0.7, suiteRevision 25)
This pull request is at predicate version v0.7, and its head carries it. The conformance corpus is at suiteRevision 25, which is 250 vectors: 55 accept, 193 reject and 2 indeterminate. The indeterminate class is a third disposition rather than a rounding of the other two, and a statement lands there when the corpus declares that more than one conformant reading is available.
aeeChainScopeis an array a machine can compare, canonically sorted and free of duplicates, of registered dimension tokens with a two-sided equality gate that closes both scope-narrowing and coarse-side pooling. The whole statement is parsed as strict I-JSON; every string literal must be a well-formed sequence of Unicode scalar values checked on the raw bytes, which now excludes the Unicode noncharacters RFC 7493 §2.1 forbids; and JSON nesting is bounded normatively at 128 with its counting rule stated (the number of open containers, the outermost brace being depth 1). An arming payload may carry a read-firstaeeBindingVersion;armedAtrequires a zero UTC offset; an out-of-rangeobservationRefsindex is a fault on any row; duplicateattackIdrows are malformed; the single-subject requirement holds on a statement of any basis; and the three coverage sets are a disjoint partition. Every interpretation decision is locked by a forcing vector under a CI-checked registry, and the spec says plainly that a reading no vector exercises is untested, not confirmed. Working in a different language from the published text, they have run this suite at six of its revisions, and not at the rest: there is no posted run for suiteRevisions 4, 7 through 21, or 23 through 25, so this pull request publishes no independent score at any of those. What they have posted is 125/125 blind at revision 1, 138/138 spec-diff-led at revision 2, 140/140 at revision 3 as a first run by an unchanged build whose reason-map rule predated the two new vectors, 149/149 at revision 5, 153/153 at suiteRevision 6 (2026-07-28, aee-checker#4), and at v0.7 and suiteRevision 22 a blind first run of 179/232 followed by a directed run of 232/232. The blind number is the informative one and it is not flattering. They partition the 53 first-run mismatches by the message their checker emitted: 42 on one run-binding derivation, 7 returning valid with no reason, and 4 answering pass_indirect where the corpus expects pass. That partition is by message and not by cause, and they state in their own report that attributing a recovery to a particular fix would need a bisection against the blind build, which is unrecoverable, so no causal residual follows from it. They also record that the blind build was never committed on its own, so the 179 is not independently reproducible and their index carries an explicit null digest saying so. Their run records and report for this revision are pinned at reports/v0.7-RUN.md, and our own implementation notes are in docs/IMPLEMENTATION-REPORT.md.A protobuf definition ships with this revision, transport/codegen-only: its JSON output is never re-canonicalized for signing (proto3 ProtoJSON is not RFC 8785), since DSSE signs the body bytes verbatim.
Edited 2026-08-20 to correct the corpus counts: suiteRevision 25 is 250 vectors, 55 accept and 193 reject, and this description carried suiteRevision 24's 248, 54 and 192 against revision 25's name.