Skip to content

Commit a0b5ca9

Browse files
authored
Merge pull request #71 from Taz33m/codex/qualification-evidence-scaffold
Add captured two-run qualification evidence
2 parents d7e20a3 + e39a010 commit a0b5ca9

10 files changed

Lines changed: 1448 additions & 5 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,12 @@ The project follows [Keep a Changelog](https://keepachangelog.com/) conventions.
66

77
## Unreleased
88

9+
- Added a captured two-run qualification workflow. Candidate adapters can now
10+
report optional revision and snapshot metadata, candidate-facing commands can
11+
enforce a task-pinned identity before the first event, and `evidence-init` /
12+
`evidence-verify` prepare clean roots and emit a grading-ready manifest only
13+
for unchanged, complete, canonical, byte-identical qualification bundles.
14+
915
## 0.6.0 - 2026-07-27
1016

1117
- Fixed conformance CLI exit classification across single runs, suites,

‎README.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -189,6 +189,15 @@ these are separate version boundaries. `fifo-limit-v1` qualification does not
189189
run STP or pro-rata cases, so unsupported features stay explicit without
190190
becoming false failures.
191191

192+
For release evidence, do not rely on one successful workspace. Start a captured
193+
pair with `tracebook-conformance evidence-init`, run the canonical qualification
194+
once in each generated clean root with the plan's pinned candidate name,
195+
revision, and snapshot, then run `tracebook-conformance evidence-verify`. The
196+
verifier emits a compact `evidence-manifest.json` only when both source trees
197+
remain unchanged and both qualification bundles agree on terminal result,
198+
candidate metadata, deterministic IDs, counts, coverage, and every artifact
199+
byte. See the [captured evidence workflow](https://github.com/Taz33m/tracebook/blob/main/docs/conformance.md#captured-two-run-evidence).
200+
192201
[Read the research-grounded roadmap and adoption experiment](https://github.com/Taz33m/tracebook/blob/main/docs/research-roadmap.md).
193202

194203
## Architecture

‎docs/api-stability.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -30,6 +30,11 @@
3030
normalization, or hashing requires an explicit protocol/schema version
3131
decision. `suite_hash` binds case configuration and fixture identity; an
3232
intentional suite edit must update it.
33+
- Protocol-v1 engine metadata may add the optional `revision` and `snapshot_id`
34+
identity fields. Ordinary adapters may omit them; a command using the three
35+
task-pinned `--candidate-*` flags requires an exact match before any event is
36+
sent. Evidence plan schema v1 and evidence manifest schema v1 are public
37+
artifacts for the canonical captured two-run qualification workflow.
3338
- Campaign artifact schema version 1 is public. Campaign generator version 2,
3439
the built-in versioned profile definitions, seed derivation, and trace hashes
3540
are reproducibility contracts. An intentional generation change requires a

‎docs/commands.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -132,6 +132,27 @@ automatically minimized failure. Use this command for a first integration or a
132132
profile-level CI claim; use `suite` when intentionally comparing every broader
133133
fixed semantic surface.
134134

135+
Prepare and verify a release-grade two-run evidence pair:
136+
137+
```bash
138+
tracebook-conformance evidence-init /path/to/candidate \
139+
--workspace .tracebook/release-evidence \
140+
--candidate-name owner/repository \
141+
--candidate-revision REVISION
142+
143+
# Run the canonical qualification from evidence-plan.json in both generated
144+
# roots, with all three --candidate-* identity flags, then:
145+
tracebook-conformance evidence-verify \
146+
.tracebook/release-evidence/evidence-plan.json
147+
```
148+
149+
`evidence-init` refuses an existing workspace and creates independent candidate,
150+
adapter, build, cache, and qualification paths for `run-1` and `run-2`.
151+
`evidence-verify` writes `evidence-manifest.json` only after both canonical
152+
bundles are qualified, byte-identical, and bound to the unchanged pinned
153+
candidate. See [Captured Two-Run Evidence](conformance.md#captured-two-run-evidence)
154+
for the exact qualification command and trust boundary.
155+
135156
Copy and run the standard suite:
136157

137158
```bash

‎docs/conformance.md‎

Lines changed: 76 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -135,6 +135,64 @@ not exchange certification and does not imply support for unselected suite
135135
cases. Selection version 1 is part of the artifact identity and will not change
136136
silently.
137137

138+
## Captured Two-Run Evidence
139+
140+
Use the evidence workflow when a qualification result will support a release,
141+
benchmark, or external claim. It pins the candidate before adapter readiness and
142+
requires the same canonical qualification from two prepared roots.
143+
144+
```bash
145+
REVISION=$(git -C /path/to/candidate rev-parse HEAD)
146+
WORKSPACE=$PWD/.tracebook/release-evidence
147+
148+
tracebook-conformance evidence-init /path/to/candidate \
149+
--workspace "$WORKSPACE" \
150+
--candidate-name owner/repository \
151+
--candidate-revision "$REVISION"
152+
```
153+
154+
`evidence-init` strips `.git` from the captured source identity, copies that
155+
exact tree into `runs/run-1/candidate` and `runs/run-2/candidate`, and creates a
156+
separate empty `adapter`, `build`, and `cache` directory for each run. It writes
157+
`evidence-plan.json` last and refuses an existing or source-nested workspace.
158+
Use each run's paths for that run's adapter and build; do not share generated
159+
files between the two runs.
160+
161+
The command prints the candidate snapshot ID. The adapter's `ready` frame must
162+
report that ID and the task-pinned name and revision. Run the plan's canonical
163+
contract in each root, changing only the run number and adapter command:
164+
165+
```bash
166+
SNAPSHOT=$(python -c \
167+
'import json,sys; print(json.load(open(sys.argv[1]))["candidate"]["snapshot_id"])' \
168+
"$WORKSPACE/evidence-plan.json")
169+
170+
tracebook-conformance qualify \
171+
--profile fifo-limit-v1 --suite-version v2 \
172+
--seed 42 --traces 25 --events-per-trace 200 --max-minimize-runs 100 \
173+
--candidate-name owner/repository \
174+
--candidate-revision "$REVISION" \
175+
--candidate-snapshot "$SNAPSHOT" \
176+
--candidate-cmd "$WORKSPACE/runs/run-1/adapter/launch" \
177+
--output-dir "$WORKSPACE/runs/run-1/qualification"
178+
```
179+
180+
Repeat for `run-2`, using only its candidate, adapter, build, cache, and
181+
qualification paths. Then verify the pair:
182+
183+
```bash
184+
tracebook-conformance evidence-verify "$WORKSPACE/evidence-plan.json"
185+
```
186+
187+
Verification fails closed if either captured tree changed; either bundle is
188+
missing, noncanonical, incomplete, or unqualified; candidate identity differs;
189+
or terminal result, deterministic IDs, counts, semantic coverage, or artifact
190+
bytes differ between runs. Success writes one exclusive
191+
`evidence-manifest.json` at the workspace root. The workflow prepares and checks
192+
the filesystem contract, but it does not sandbox a compiler or prove that an
193+
adapter avoided undeclared external caches; CI or a disposable host remains the
194+
strongest execution boundary.
195+
138196
## Differential Campaigns
139197

140198
Campaigns generate stateful traces, compare them one at a time, and stop at the
@@ -346,9 +404,15 @@ Host to candidate, once:
346404
Candidate to host:
347405

348406
```json
349-
{"type":"ready","protocol":"tracebook.conformance","protocol_version":1,"engine":{"name":"my-engine","version":"1.4.2","language":"Rust"}}
407+
{"type":"ready","protocol":"tracebook.conformance","protocol_version":1,"engine":{"name":"owner/repository","version":"1.4.2","language":"Rust","revision":"8f31c2a","snapshot_id":"sha256:..."}}
350408
```
351409

410+
`revision` and `snapshot_id` are optional for ordinary comparisons. When
411+
`--candidate-name`, `--candidate-revision`, and `--candidate-snapshot` are
412+
supplied, all three flags are required together and the `ready` metadata must
413+
match before Tracebook sends the first event. Captured two-run evidence always
414+
uses the pinned form.
415+
352416
Host to candidate for each event:
353417

354418
```json
@@ -429,7 +493,13 @@ from tracebook.conformance import EngineMetadata, serve_stdio
429493

430494
class MyAdapter:
431495
def __init__(self, config):
432-
self.metadata = EngineMetadata("my-engine", "1.0", "Python")
496+
self.metadata = EngineMetadata(
497+
"owner/repository",
498+
"1.0",
499+
"Python",
500+
revision="8f31c2a",
501+
snapshot_id="sha256:...",
502+
)
433503
self.config = config
434504

435505
def apply(self, event, index):
@@ -447,6 +517,10 @@ class MyAdapter:
447517
raise SystemExit(serve_stdio(MyAdapter))
448518
```
449519

520+
The last two fields are needed only when the invoking command pins candidate
521+
identity. In a captured evidence run, populate them from the immutable task or
522+
`evidence-plan.json`, not by inspecting the host after the adapter starts.
523+
450524
[`examples/conformance_adapter.py`](../examples/conformance_adapter.py) is a
451525
runnable reference. Non-Python adapters implement the same frames directly.
452526

‎src/tracebook/conformance/cli.py‎

Lines changed: 82 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -20,11 +20,12 @@
2020
write_campaign_corpus,
2121
)
2222
from .compare import run_conformance
23+
from .evidence import prepare_evidence_workspace, write_evidence_manifest
2324
from .external import AdapterProtocolError, ExternalProcessAdapterFactory
2425
from .exit_codes import exit_code_for_artifact
2526
from .junit import write_junit
2627
from .minimize import minimize_failing_trace
27-
from .model import ConformanceConfig, ConformanceError
28+
from .model import ConformanceConfig, ConformanceError, PinnedCandidateIdentity
2829
from .qualification import _QualificationOutputReservation, run_qualification
2930
from .reproduce import (
3031
discover_failure_metadata,
@@ -54,6 +55,18 @@ def _add_config_arguments(parser: argparse.ArgumentParser) -> None:
5455

5556
def _add_candidate_arguments(parser: argparse.ArgumentParser) -> None:
5657
parser.add_argument("--timeout", type=float, default=5.0)
58+
parser.add_argument(
59+
"--candidate-name",
60+
help="Task-pinned candidate name expected in the ready frame.",
61+
)
62+
parser.add_argument(
63+
"--candidate-revision",
64+
help="Task-pinned revision expected in the ready frame.",
65+
)
66+
parser.add_argument(
67+
"--candidate-snapshot",
68+
help="Task-pinned snapshot ID expected in the ready frame.",
69+
)
5770
candidate = parser.add_mutually_exclusive_group(required=True)
5871
candidate.add_argument(
5972
"--candidate-cmd",
@@ -163,6 +176,23 @@ def _build_parser() -> argparse.ArgumentParser:
163176
reproduce.add_argument("--junit-output")
164177
_add_config_arguments(reproduce)
165178
_add_candidate_arguments(reproduce)
179+
180+
evidence_init = commands.add_parser(
181+
"evidence-init",
182+
help="Create two clean roots for a captured qualification pair.",
183+
)
184+
evidence_init.add_argument("candidate_source")
185+
evidence_init.add_argument("--workspace", required=True)
186+
evidence_init.add_argument("--candidate-name", required=True)
187+
evidence_init.add_argument("--candidate-revision", required=True)
188+
evidence_init.add_argument("--expected-snapshot")
189+
190+
evidence_verify = commands.add_parser(
191+
"evidence-verify",
192+
help="Verify two captured qualification bundles and write one manifest.",
193+
)
194+
evidence_verify.add_argument("plan")
195+
evidence_verify.add_argument("--output")
166196
return parser
167197

168198

@@ -174,7 +204,32 @@ def _candidate_factory(args) -> ExternalProcessAdapterFactory:
174204
command = command[1:]
175205
if not command:
176206
raise ConformanceError("--candidate requires a command")
177-
return ExternalProcessAdapterFactory(command, timeout_seconds=args.timeout)
207+
identity_values = (
208+
args.candidate_name,
209+
args.candidate_revision,
210+
args.candidate_snapshot,
211+
)
212+
if any(value is not None for value in identity_values) and not all(
213+
value is not None for value in identity_values
214+
):
215+
raise ConformanceError(
216+
"--candidate-name, --candidate-revision, and --candidate-snapshot "
217+
"must be provided together"
218+
)
219+
expected_identity = (
220+
PinnedCandidateIdentity(
221+
name=args.candidate_name,
222+
revision=args.candidate_revision,
223+
snapshot_id=args.candidate_snapshot,
224+
)
225+
if all(value is not None for value in identity_values)
226+
else None
227+
)
228+
return ExternalProcessAdapterFactory(
229+
command,
230+
timeout_seconds=args.timeout,
231+
expected_identity=expected_identity,
232+
)
178233

179234

180235
def _config(args) -> ConformanceConfig:
@@ -253,6 +308,31 @@ def main(argv: Optional[List[str]] = None) -> int:
253308
print(f"Suite id: {suite.suite_id}")
254309
print(f"Cases: {len(suite.cases)}")
255310
return 0
311+
if args.command == "evidence-init":
312+
plan_path = prepare_evidence_workspace(
313+
args.candidate_source,
314+
args.workspace,
315+
candidate_name=args.candidate_name,
316+
candidate_revision=args.candidate_revision,
317+
expected_snapshot=args.expected_snapshot,
318+
)
319+
plan = json.loads(plan_path.read_text(encoding="utf-8"))
320+
print(f"Evidence plan written: {plan_path}")
321+
print(f"Candidate snapshot: {plan['candidate']['snapshot_id']}")
322+
for run in plan["runs"]:
323+
print(
324+
f"{run['run_id']}: candidate={run['candidate_root']} "
325+
f"qualification={run['qualification_dir']}"
326+
)
327+
return 0
328+
if args.command == "evidence-verify":
329+
manifest_path = write_evidence_manifest(args.plan, args.output)
330+
manifest = json.loads(manifest_path.read_text(encoding="utf-8"))
331+
print(f"Evidence manifest written: {manifest_path}")
332+
print(f"Manifest: {manifest['manifest_id']}")
333+
print(f"Qualification: {manifest['runs'][0]['qualification_id']}")
334+
print("Evidence pair: PASS")
335+
return 0
256336
if args.command == "run":
257337
_require_distinct_paths(args.events, args.output, args.junit_output)
258338
single_report = run_conformance(

0 commit comments

Comments
 (0)