report.json is the machine-readable evidence sidecar for one reduction. It is
written next to the exported payload at OUTPUT.repomin/report.json; keeping
it outside the payload means report writes cannot change a tree that already
passed the oracle.
The current top-level schema version is 1. Consumers should reject an
unsupported schema_version, tolerate additional fields within a supported
version, and never infer code correctness from a passing oracle.
| Field | Meaning |
|---|---|
schema_version |
Integer report format version. Current value: 1. |
repomin_version |
ReproMin version that generated the report. Optional in legacy reports. |
command |
Exact reproduction command passed to the runner. |
failure_match |
Configured output regular expression, or null for process/exit-code modes. |
failure_spec |
Exact match, exit-code, and signature-mode flags used for replay. Optional in legacy reports. |
baseline_exit_code |
Return code observed during baseline validation. |
final_exit_code |
Return code observed during final validation of the accepted tree. |
source |
File/byte counts for the copied source tree before reduction. |
output |
File/byte counts for the exported payload, excluding the sidecar. |
attempts |
Logical candidate attempts, including no-ops and cache uses. |
accepted_mutations |
Number of promoted candidate mutations. |
cache_hits |
Session-local content-cache uses. These are not oracle executions. |
execution |
Runner, sampling, ignore-rule, and resource configuration. |
phase_statistics |
Per-phase accounting and oracle sample usage. |
holdout_certification |
Optional fresh-sample certification of the exported artifact. |
events |
Ordered human-readable reduction events and their oracle evidence. |
java_exception_signature |
Present only with --java-exception. |
python_exception_signature |
Present only with --python-exception. |
process_failure_signature |
Present only with --process-failure. |
source and output contain files and bytes. Output counts deliberately
exclude report.json and REPOMIN.md. New reports also store
output.tree_sha256 and output.tree_fingerprint_policy for every exported
payload, independently of optional holdout certification. They also store
output.tree_content_sha256 with the tree-content-sha256-v1 policy. The
complete fingerprint includes filesystem metadata and is authoritative for a
local export; the content fingerprint covers paths, entry kinds, file contents,
and symlink targets so archive transports that rewrite mtimes can still be
checked. A consumer should label that case as content-only verification.
failure_spec is an additive schema-v1 object that preserves the exact oracle
configuration needed by replay: match, optional exit_code, and the boolean
java_exception, python_exception, and process_failure modes. At most one
signature mode can be true, process-failure mode cannot also configure an exit
code, and the stored match must equal the legacy top-level failure_match.
Java, Python, and process signature objects remain top-level fields for
backward compatibility. When failure_spec selects a signature mode, exactly
the corresponding recorded signature must be present. Replay pins this
identity; it never learns a replacement signature from current output.
The execution object records the boundary in which commands were sampled.
Important fields include:
backend:hostordocker.jobs: maximum candidate concurrency.cache_enabled,cache_hits, andresumed.baseline_runs,candidate_runs,final_runs, and their pass counts.confidence,min_baseline_rate,min_candidate_rate, and the sampling policy identifiers.reduction_strategy: the reducer strategy identity used for this report and persistent-session compatibility.ignored_names,ignored_paths,gitignore_files,keep_paths, andtext_files: input-selection controls applied before reduction.environment_namesandenvironment_sha256: names and a digest of explicit environment values. Values are intentionally never recorded.timeout_seconds: configured timeout for each reproduction command.semantic_reducer,semantic_model,semantic_endpoint, andsemantic_timeout: opt-in semantic backend provenance. The timeout is the positive HTTP request timeout in seconds, ornullwhen that backend is disabled. Legacy reports may omit it.budget_exhausted: boolean indicating whether an optional reduction budget stopped the search before the normal fixed-point condition.
Docker reports additionally contain the image reference, resolved immutable image ID, network policy, and configured resource limits when applicable. These fields describe the execution boundary; they do not make Docker a complete security sandbox.
phase_statistics.phases contains one object per reduction phase. Each phase
tracks attempts, no-ops, rejected/accepted/superseded/aborted candidates,
oracle sample uses, actual oracle samples, cache hits, and samples saved by
early stopping.
For complete reports, consumers can check both accounting identities:
attempts = no_op + rejected + accepted + superseded + aborted
oracle_sample_uses = oracle_samples + cache_hits
coverage is partial when a legacy or interrupted session cannot provide a
complete phase history. Missing historical data must not be reconstructed from
the aggregate counters.
holdout_certification.status is not_requested, certified, rejected, or
an interrupted/aborted status. When certification is enabled, its samples are
fresh fixed-size runs against the frozen exported payload. They are separate
from baseline, candidate, and ordinary final-validation samples.
The report records the planned/completed sample counts, passes, exact lower
bound, exact p-value, resource/timeout veto counts, artifact fingerprint, and
the holdout policy identifier. Sample index values are one-based and
contiguous through completed_runs. When present, each sample's outcome must
agree with its acceptance, timeout, and resource-exhaustion flags; interrupted
samples carry no execution evidence. A certified lower bound is a statistical
claim about oracle pass probability under fresh iid samples in the recorded
environment. It is not a proof of correctness, compatibility, or production
reliability.
When the aggregate timeout, resource-exhaustion, or interruption counters are
present alongside complete sample fields, they must equal the corresponding
sample counts and cannot exceed completed_runs.
Terminal holdout statistics are an all-or-none group when present. Modern
reports (holdout_certification.schema_version: 1) with status certified
must include the complete terminal group and must have completed every planned
run. The observed_rate must equal passes / planned_runs, confidence must be
in (0, 1), and the exact gate result must be boolean. When the optional
ordinary_failures aggregate is present, it must equal the number of samples
whose outcome is failed.
Each events entry records the phase, description, duration, oracle pass/runs,
rate and lower-bound evidence, and (when applicable) candidate family
confidence and early-acceptance state. Event order is significant for audit
and resume diagnostics. Optional rates and bounds are finite probabilities in
the inclusive [0, 1] range; oracle_rate must agree with
oracle_passes / oracle_runs. Candidate family index, confidence, and alpha
are an all-or-none group when present, and the early-acceptance flag is
boolean when present.
Signature objects preserve identity beyond a broad output match. Java and Python signatures include exception class, message, and normalized frames. Process signatures distinguish POSIX signals, Windows statuses, and ordinary exit codes. Timeout and resource-exhaustion outcomes are never treated as a matching failure signature.
- Verify
schema_version, output file/byte counts, and the exported payload fingerprint before trusting an artifact. A content-only match means transport metadata may have changed. - Check
execution.backend, Docker identity/policy, environment names, and the reproduction command before sharing the sidecar. - Treat
failure_matchand signatures as the configured oracle contract, not as an explanation of every possible failure mode. - Keep
report.jsonandREPOMIN.mdbeside the payload; do not copy either file into the tree when independently rerunning the command.
The bundled validator checks these structural rules without executing the reproduction command:
repomin report validate OUTPUT.repomin/report.json --payload OUTPUTIt returns exit code 2 for malformed JSON, unsupported schema versions,
inconsistent phase/holdout accounting, unsafe payload entries, size drift, or
a payload fingerprint mismatch. Legacy reports without an output or holdout
fingerprint still receive safe-tree and file/byte-count validation.
Add --json when a CI step needs a compact result. The
current result has summary_schema_version: 2 and includes valid,
schema_version,
holdout_status, the resolved report path, an optional resolved payload path,
repomin_version, backend, the
privacy-safe oracle_mode, source/output file and byte counts, removed
file/byte counts, file/byte retention ratios, attempts,
accepted_mutations, cache_hits, and budget_exhausted. Holdout run/pass
counts are included without including command output,
environment names, or environment values. The version is null for
legacy reports that predate version provenance. Ratios are null when the
source denominator is zero; otherwise they are descriptive fractions rounded
to six decimal places. A negative removal count is possible for a report whose
recorded output is larger than its source and should be read as a size change,
not a correctness signal. These fields are descriptive metadata copied or
derived from the validated report; they do not add a new correctness claim.
payload_fingerprint_verified distinguishes a cryptographic tree match from
count-only validation of a legacy report, and payload_fingerprint_mode is
exact, content, or unavailable when --payload is supplied.
Because the JSON contains local paths for diagnostics, redact those fields
before posting it publicly. Use --format markdown for the fixed, path-free
shareable field set described below.
Summary schema version 1 (used by v0.1.0.dev7) included an
environment_names_count field. Version 2 removes that field so shareable
summaries contain no environment metadata; consumers should branch on
summary_schema_version instead of assuming fields are stable across releases.
Malformed or overlong repomin_version provenance is represented as null in
the summary rather than copied into shareable output.
For a human-readable, shareable version of the same safe fields, request the Markdown exporter:
repomin report validate OUTPUT.repomin/report.json \
--payload OUTPUT --format markdownThe exporter is deterministic and uses a fixed whitelist: summary/report
schema and version, execution backend, oracle type, source/output sizes,
removal and retention figures, reduction attempt/mutation counts, holdout
status/counts, and payload-fingerprint status. Every value is escaped for a
Markdown table. It never renders the report path, payload path, reproduction
command, match expression, logs, environment names or values, signatures, or
other report fields. A missing payload is represented as n/a for fingerprint
fields; it does not trigger command execution. Invalid JSON, unsupported schema,
inconsistent evidence, or a mismatched payload returns exit code 2 and emits
no Markdown summary.
After reviewing the unsigned command in a report, consumers can also run a fresh-copy replay:
repomin report replay OUTPUT.repomin/report.json --payload OUTPUT --yesReplay is a new current-environment observation. It does not upgrade the original report or create a statistical certificate.
repomin report compare REPORT.json REPORT.json ... validates at least two
reports and returns a separate comparison document. The command accepts the
paths in the order supplied; it does not sort by filename, version, or time.
It never reads a payload, executes a recorded command, or accesses the network.
The top-level comparison fields are:
| Field | Meaning |
|---|---|
comparison_schema_version |
Integer comparison format version. Current value: 1. |
descriptive_only |
Always true; the output is evidence metadata, not a correctness or performance claim. |
run_count |
Number of validated reports in the comparison. |
runs |
Ordered snapshots using the fixed allow-list below. |
deltas |
Adjacent snapshot differences, calculated as next minus previous. |
context_warnings |
Deterministic warnings for changes that make snapshots harder to compare directly. |
Each runs entry contains only index, display-only label, safe
repomin_version provenance (or null for legacy/unusable provenance),
backend, an enumerated oracle_mode, source/output file and byte counts,
file/byte retention ratios, attempts, accepted_mutations, cache_hits,
budget_exhausted, holdout status and planned/completed/pass counts, and
phase_coverage. Ratios are rounded to six decimal places and are null when
the source denominator cannot be represented safely as a finite ratio. Labels
must be unique, short ASCII identifiers and affect display only.
Each deltas entry identifies the adjacent from_* and to_* rows, provides a
numeric_deltas object for the snapshot numeric fields, and lists changed
categorical fields in changed_fields. A delta is descriptive; it does not
attribute a change to a particular version, option, backend, or mutation.
Integer counters remain exact when their adjacent difference is within the
comparison output bound; an excessively large difference is represented as
null rather than expanding a shareable document without limit.
Warnings cover unavailable or changed version provenance, source size, input
selection/exclusion controls, backend, jobs, timeout/cache and budget settings,
semantic/container/environment and working-directory context, oracle mode or
identity, candidate/baseline sampling configuration, reduction strategy,
holdout controls/policy/status, and phase definitions/coverage, including
partial coverage or unavailable ratios. Private fields are compared only by
an internal opaque digest and are never copied into the result. The comparison
intentionally excludes paths, commands, match expressions, logs, environment
names/values, signatures, fingerprints, semantic endpoints, and phase timing.
At most 32 reports can be compared in one invocation.
Use benchmarks/compare.py or a dedicated benchmark system for performance
history. Consumers should branch on comparison_schema_version, tolerate only
documented additive fields, and preserve the descriptive_only boundary when
sharing results.
The architecture document explains the statistical contracts and reducer invariants behind these fields. See ARCHITECTURE.md and SECURITY.md before processing untrusted repositories.