Skip to content

Latest commit

 

History

History
1005 lines (1002 loc) · 69.9 KB

File metadata and controls

1005 lines (1002 loc) · 69.9 KB

Glossary

Terms describe the harness unless explicitly marked as UI design concepts. See current state for implementation coverage and the reference for command and delivery contracts.

  • Accepted band — the observed word range of accepted rewrites used to design the brief-reply cases. Their word-band checks declare a ceiling rather than a minimum. This is case-specific evidence, not a general rule for reply length.

  • Artifact — durable output of a stage: a spec or plan document, a backlog card update, or commits.

  • Assessment — the record one regrade writes: the attempt it read, the digest of each evidence body it read, the grading definition digest that produced it, one result per declared check, and a state grade where the case declares a scorer. It carries an overall outcome only when every declared check was graded, because a verdict over the checks that had evidence would read as a verdict over the whole definition. It is filed beside the attempt and never rewrites the attempt record.

  • Attempt — one execution of a case's unit of work: a stage at a checkpoint (the original run's stage result or any replay) or one session of a session case. The unit a comparison presents.

  • Attempt record — the strict record a session attempt writes, including case, lineage, model, declared corpus digests, prompt, transcript evidence, checks, and provider metrics when available. A completed reply, no reply, and an invocation failure have distinct outcomes. The record may preserve a supplied provider context evidence bundle and its normalized projection. Older records may omit snapshot provenance and context evidence.

  • Attempt region — the part of a resumed session's transcript the attempt itself produced: the 1-based physical lines after the prefix line count the attempt record carries. The lines at or before that count are the inherited starting context, and a transcript whose record carries no prefix count is boundary unknown throughout, the same three regions the context history projects (see reference). Only attempt-region requests count toward an attempt's totals: on a measured resumed attempt the region's two requests match the provider envelope in all four usage categories, while the whole transcript sums 90743 output tokens against the envelope's 151, because the rest belongs to the session that was resumed. A transcript whose boundary is unknown, and an attempt that saved no transcript, both report totals unavailable, since zero and a whole-transcript sum would each assert something the record does not settle.

  • Attempt directory — the fresh temporary directory the harness creates and owns for one session attempt, seeded from the case's fixture tree when it declares one. A session attempt never runs in a live repository, and the directory's real path is what names the attempt's project slug.

  • Attempt state evidence — the copy of the attempt directory preserved beside the transcript before cleanup removes it, holding the files and git state the session left: dirty tracked files, untracked files, ignored files, and the commit history, with .git stored as dot-git because git refuses to commit a nested repository. The corpus overlay is excluded, being an input the record already digests per file. It is what a state check reads, and each grade reads its own restored copy, so a grader that writes changes neither the evidence nor a later pass's input. An attempt whose provider call failed preserves none.

  • Attribution (comparison) — the claim permitted by corpus differences between two named comparison arms after their recorded executed-corpus entries are normalized and deduplicated by corpus layout path. No differing path means the corpora are identical; exactly one means a movement can be attributed to that path; more than one refuses attribution and names every differing path.

  • Calibration — the step that validates a Judge result against a human review and turns findings into rubric or instruction changes. It reads the frozen evidence a run recorded and the corpus files as they stand now, never the live target, so it can run long after the target was restored. It is one function of the review, the frozen evidence, and the current rubrics and instructions, whether a paused run calls it in a retry loop or the calibrate command calls it once. The browser's Judge calibration screen is a different reading, grade agreement and Judge drift over operator grades, and changes no grade.

  • Checkpoint — frozen input state that can start a stage, containing target SHA, workflow state, artifacts, and lineage. A run records an initial checkpoint after task setup and further checkpoints after accepted stages.

  • Check-integrity file — a target-relative file declared by the pipeline whose presence and bytes are frozen at baseline and compared after delivery.

  • Check kind — one deterministic assertion a session case may declare, the discriminator of a check: word-band, forbidden-text, and forbidden-pattern read the reply, tool-calls and files-read read the transcript. A state check grades the files and git state the session left; it is declared beside the check list rather than inside it, and carries no kind of its own. A kind states what it needs and what it reports; the case supplies the values it compares against, so no literal a case could differ on lives in the check.

  • Check list — the ordered deterministic checks a session case declares as its judge. It is evaluated over the reply and the transcript, needs no provider call, and its rep outcome is successful when and only when every check passes.

  • Clean record — a recorded result judged against the corpus under test with no staleness cause and a recorded version distance of zero. Run history marks it ✓ clean. A record whose version distance is not recorded is never clean, whatever its causes. The design's Clean corpus only filter names these. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-249).

  • Confirmation run — an explicitly requested group of at least two reps over one frozen input set, used by the outer loop to produce a score; defaults to five reps. The UI's design calls this a group and counts it in plural attempts (group · 6 attempts, the ×1/×3/×6/×12 replay control, the Cases screen's Run group action); the word in code, records, and this glossary stays confirmation run (see UI vocabulary).

  • Command — one named verb of the rehearse executable (run, replay, compare, review, calibrate, list, show, stale, case list, case show, case capture), declaring its own flags with their defaults, environment fallbacks, and help lines as data. A name is one or two tokens; the longer declared name wins over a prefix of it. The declaration is the single source of the flag's name in help, parsing, and documentation, and a command with no declared flag prints no flag section.

  • Command record — the validated evidence a command writes. Depending on mode, run returns a run artifact, session attempt, or confirmation report. --json selects record bytes over the record path; stdout carries one or the other and no progress, which goes to stderr. show --json reads the selected record directly.

  • Comparison — a deterministic report over completed stage, pipeline, or session confirmation evidence, each case carrying baseline, candidate, and control arms. A pipeline comparison needs at least two benchmark cases; a stage comparison of one replayed checkpoint, or a session comparison, may cover one. It starts no paid sessions. Session comparisons read the frozen case's checks and each recorded attempt, and do not synthesize a pipeline final outcome. A current report keeps each source rep's ordered quality outcomes beside its identity, so a reader can associate an ordinal with its grades and non-judged statuses without reopening source records. Older reports remain readable with the fields they were written with, and a measurement a version predates reads as unavailable rather than as zero.

  • Word count — how many whitespace-separated words an attempt's output holds, which is how verbosity gets caught. A session attempt's output is its reply. A stage attempt's or replay's output is its artifact, the text the stage judge read, or, for a stage that wrote no artifact such as a delivery stage, the worker's final reply, never the diff, since code length says nothing about verbosity. A pipeline attempt counts its last declared stage, and one stopped before it reads unavailable. The count is read from the output the record holds. An attempt with no output reads unavailable with its reason, never zero words.

  • Attempt pairs (comparison): each arm's recorded attempts listed side by side, each with its outcome, the blockers that fired on it and its words. Despite the name, no attempt is paired with another: nothing recorded ties one arm's attempt to another's, so the list claims no per-pair change.

  • Attempt combinations (comparison): how one arm's attempts compare with another's over every pairing of one attempt from each, counted higher, equal and lower: n x m for arms of n and m attempts on a pass/fail measure, and over graded attempts only on a graded one. Every combination counts because nothing recorded ties one arm's attempt to another's.

  • Comparison extension: a new comparison holding a saved comparison's attempts and n more in every arm, run at the cost stated before it starts. It names the comparison it extends, which is kept but no longer listed.

  • Average words (comparison arm) — the mean word count over an arm's attempts that have one, served beside how many attempts it counted and how many the arm holds. A report written before word counts reads unavailable.

  • Comparison arm — one role in a comparison: baseline, candidate, or the mandatory control. An arm uses the same corpus snapshot across every benchmark case; a session control may have an empty declared corpus.

  • Baseline arm, arm A and arm B: the names the comparison screen and compare attempts give the three comparison arms. The baseline arm is the control, arm A is the baseline role, the corpus before the edit, and arm B is the candidate role, the corpus after it. A baseline arm compare attempts derived is arm A's corpus with the skill under test removed, one it took unchanged is arm A run again, and one a manifest supplied is described as a minimal corpus.

  • Skill under test: the one corpus unit that differs between arms A and B of a compare attempts comparison, which a derived baseline arm removes to show the corpus does work at all. A manifest comparison names none. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-257, doc-229).

  • What moved (comparison): the per-case rows a comparison answers "did the change help, hurt or do nothing" with, in order: the overall grade, each hard blocker's firings, each quality dimension's letter span, reply length and cost per attempt. A blocker fired when its condition occurred on an attempt; its row reads each arm's firing rate by its 95% Wilson interval and names the arm that fires less only when the intervals separate. Reply length and cost are meters: each arm's mean and low-to-high range over its attempts, with the signed percent change of the means, and a higher arm named only when the ranges do not overlap and that separation would happen by rerun noise at most one time in twenty, which takes about four attempts an arm. A row where an arm recorded nothing reads unavailable, never rerun noise.

  • Quality reading (comparison): a per-case, per-arm-pair, per-measure interpretation of how far repeated attempts spread. A stage grade carries each arm's observed low-to-high letter span. A pass/fail measure, a session's checks or a pipeline's final verdict, carries each arm's 95% Wilson interval on its success rate, because an arm holding both outcomes spans the whole pass/fail scale however rarely it fails. The reading gives one verdict: the intervals overlap within rerun noise, both arms already succeed on every requested rep, or the intervals are separated and the arm that succeeds more often is named. At a few reps the interval is wide, so 3 of 3 against 0 of 4 still reads as rerun noise. It is derived from each arm's recorded reliability summary, not from the paired estimate across cases.

  • Case disagreement — the reading a multi-case comparison prints when the per-case deltas of one contrast do not share a direction: one case moved up and another moved down. It is reported beside the mean because a mean over cases describes the cases only when they agree, and deltas of +0.5 and -0.5 average to the same zero as two cases that did not move at all. Sign is the test rather than the mean's distance from its standard error, since a genuinely null result also sits near zero and is not a disagreement. Unlike the quality reading, it is derived from the paired estimate across cases.

  • Benchmark case — one frozen task with its source or checkpoint and all non-corpus inputs. It is the sampling unit a multi-case comparison pairs its arms on; a single-case stage or session comparison samples reps instead.

  • Sampling unit — the observation a comparison's uncertainty is estimated over. A multi-case comparison pairs its arms case by case and reads variation between case means. A single-case stage or session comparison samples the reps of one case, treats the arms as independent samples rather than paired, and prints the unit it used beside the estimate.

  • Minimum detectable effect: the smallest true difference a comparison of a given size and spread would detect at a stated confidence and power. A comparison that finds no difference supports "no effect" only down to this size. Rehearse does not compute it.

  • A/A check: a comparison whose arms run identical inputs, so its interval should contain zero. It tests whether the comparison's uncertainty is honest. Rehearse has no A/A mode, and its baseline derivation and stage replays refuse arms with identical inputs.

  • Positive control: a comparison against a change known to matter, such as deleting the instruction a case exercises, which a working comparison must detect before its null readings are trusted. Rehearse does not run one.

  • Case (design usage) — the UI design's phrase for a task plus the corpus, judges, and thresholds it runs under. It overlaps with this glossary's benchmark case without matching field for field: the case declaration pins task, product brief, final rubric, per-stage rubrics, pipeline, and target, but has no case-level corpus field (corpus is chosen per run by --corpus) and no case-level threshold field (a minimum grade is a run setting, not part of the case declaration) (see UI vocabulary).

  • Case declaration — the case.json, committed or not, that states a benchmark case as data: its id, kind, title, and the case-relative inputs the kind needs. It is parsed at the boundary. Case input paths stay within the case directory, while a pipeline target may point to an external repository.

  • Case directory — cases/<id>/, the one place a case's declaration and its input files live, transcript prefixes included. The directory name is the case id. Cases live in the control repository, never beside the corpus they grade.

  • Case kind — which inputs a case declares and how an attempt at it is run: pipeline, today's stage graph against a target repository, or session, one Claude session. The kind is the discriminator of the case declaration, so a case cannot carry another kind's inputs.

  • Control repository — this repository, containing the harness, case definitions, and rubrics, with local run artifacts under its ignored run directory. Its contributor instructions are separate from the corpus under evaluation.

  • Contribution — one of the three run-detail UI layouts. It grades a run's outcome on its own, from recorded evidence, then has an agent (not a deterministic computation) name a likely root-cause stage among those that ran. The agent's reading is disclosed as an opinion, never as a measurement, and is read only once the run has ended. It is not an ablation: ablation needs a rerun per node and is a separate planned feature (see UI vocabulary).

  • Step rail: the run-detail layout that shows one stage at a time: the stage list on the left with the attempts at the checkpoint the selected stage started from, and the stage's report on the right. A run opens on it, at the stage it stopped at or else the last stage with a record.

  • Record ledger: the run-detail layout that shows every stage's record top to bottom, with a note for each stage an ended run never reached.

  • Step modal: the dialog a task-graph node's in / out action opens: the instruction files and artifacts a stage started from, what came out, its judge summary, and the operations on it.

  • Context manifest — the transcript-observed instruction/context paths for a session attempt, classified as corpus or project inputs and reconciled against declarations. Observed entries are name-only; declared corpus hashes are separate evidence. A file counts when an instructions attachment names it or a Read of it is not answered with an error, a skill counts when its body arrives, and the output style counts when its attachment names it, so a refused call records nothing. A load from a .claude outside the session's corpus, or from the attempt's .claude that the harness did not install, is left out. A load does not prove the instruction was followed, and missing observations do not prove absence from context. The manifest deduplicates paths rather than retaining a load history. It is built for a session attempt; a pipeline stage is observed through context history instead, and no manifest is built for one.

  • Read manifest: what one pipeline stage, replay or session attempt declared or was observed to load, with each file's role and hash. It extends the context manifest to stage checkpoints and replays and adds judge rubrics. Each entry names its half (corpus, project or rubric), one role (global instructions, project instructions, stage skill, judge rubric, read for context), and whether it was declared, observed, or both. A corpus file's hash is the bytes the record's corpus resolved at its start, a rubric's is the frozen scorecard's rubric as parsed, not its bytes, and a project file's is the target's bytes at the record's start, absent when the file was not there. A corpus file is one loaded from where the record's corpus resolved it, and under a .claude the corpus was installed into only the files installed or resolved there at the start count. Any other file a stage, replay or rep loaded from the target's own .claude is a project file. A session attempt leaves out a load from its .claude that the harness did not install, and every record leaves out one from any other .claude. Like the context manifest, a load not observed is not proof of absence.

  • Context evidence — an optional, versioned attempt-record field containing an unchanged provider capture and the harness's normalized projection. The projection joins request usage, model, provider cost, agent parentage, instruction loads, and compactions when documented identifiers support the join. Request identities are session-scoped, and client/server aliases are retained. The projection keeps missing identifiers and conflicts explicit. Malformed records from validated hook, OTel, raw-body-reference, and coverage shapes remain explicit. Calculated request cost retains the selected rate and its frozen catalog; omission of the field means no provider bundle was supplied.

  • Context half — whether an instruction/context path or divergence is a corpus input, from the instruction files under evaluation (corpus), or a project input, from the case's fixture tree (project). Keeping the two halves distinct lets divergences name which side was expected without conflating a missing fixture file with a missing skill. A record written before entries carried a half has none, which means not recorded rather than corpus.

  • Context history — the ordered, read-only browser projection of one saved session attempt's or pipeline stage's starting context, tool events, observed deliveries, results, and evidence gaps. It is derived from the colocated transcript and is not a measurement of the provider's active context window.

  • Corpus (instruction corpus) — the instruction files under evaluation: the installed CLAUDE.md, the stage skills, the output styles, the agent definitions, and the rulebook. A case names the ones it reads in corpus layout paths (CLAUDE.md, skills/<name>/..., output-styles/<name>.md, agents/<name>.md, rulebook/<name>.md), which one resolver maps onto the install, so an edit to any of them can make a prior attempt stale.

  • Corpus version: one state of the whole corpus layout of a corpus source, every file the layout holds whether or not a stage reads it, identified by the sha256 of its canonical file list. A pipeline run measures it when it starts, a pipeline stage, replay and session attempt before their session, and a confirmation group once when it freezes its inputs. The records directory keeps each version openable. A layout that refuses hashing yields a corpus refusal in its place. Name accepted unattended as unsettled, pending the operator's confirmation (doc-157, question 6).

  • corpus@<hash>: the label of a corpus version, corpus@ and the first six hex characters of its digest. Run history, the navigation rail's corpus card and the corpus screen all show it, and a command that opens a version accepts any unambiguous prefix of the digest, with or without corpus@.

  • Cut — the 0-based line index of the first session-file record a transcript prefix drops. A cut of N keeps lines [0, N).

  • Corpus layout — the directory shape a corpus takes once resolved, and the only shape the harness reads: CLAUDE.md, skills/<name>/, output-styles/<name>.md, agents/<name>.md, and rulebook/<name>.md under one root. A corpus layout path names a file within it. Every corpus source resolves to this layout, so the code that hashes and installs a corpus never learns where the bytes came from.

  • Corpus layout path — how a case names a corpus file, independent of where the corpus is installed: CLAUDE.md, output-styles/<name>.md, agents/<name>.md, rulebook/<name>.md, or skills/<name>/.... One resolver maps a layout path onto the selected corpus root. Missing files are refused, although a debug command may already have paid for a model probe.

  • Corpus snapshot — a recorded set of corpus inputs and their provenance. Session execution copies the corpus files a case declares, and confirmation copies them once for the whole group. A live session debug snapshot keeps a pointer to the installed files only when the case declares no corpus file. Stage snapshots capture the instructions and supporting files needed for replay. Hashing an input is distinct from delivering it to the provider.

  • Corpus overlay — the declared styles, agent definitions, rulebook files, skills, and CLAUDE.md written under a session attempt's .claude/ directory. A declared skill brings its whole directory. Stage replay uses a separate snapshot installation path. See the reference's corpus support matrix.

  • Corpus refusal — a named statement that a corpus path selected for reading cannot supply the directory entries or file bytes its role claims, carrying the layout path and the reason: it resolves outside the permitted extent, its link target is missing, its link never resolves, it cannot be read, or it has the wrong file type. A refusal is data a report carries, not a failure of the report: the reading surface names the entry, omits its layout directory's files, and withholds the corpus digest, so an unidentifiable corpus is never served as an identified one.

  • Corpus snapshot origin — where a snapshot's bytes were read from, recorded beside them and persisted in the attempt record: the live install, or the directory the source named. What produced that directory is not recorded, because the harness never learns it.

  • Corpus source — where an attempt's corpus bytes come from, named by --corpus: a directory already in corpus layout, and nothing else. A corpus that lives somewhere else is rendered to a directory with whatever tool owns it, outside Rehearse, and that directory is passed. Absent --corpus the source is the directory linked in the settings, or the live install when none is linked. Whether bytes are copied and delivered depends on the execution mode, as described in the reference support matrix. The live source's permitted extent is its install root and one backing tree declared outside the corpus. A link may resolve within either tree; the corpus's own links cannot declare another permitted tree.

  • Linked corpus: the corpus source stored in the settings, which every command measures when none is named with --corpus. It is either a directory in corpus layout or, when nothing is linked, the live install. A pipeline run reads only the live install and refuses while a directory is linked.

  • Corpus tier — stage-local (a skill; testable in stage mode) or global (CLAUDE.md, doctrine; validated only end-to-end).

  • Corpus variant — one corpus a comparison arm runs against, identified by the snapshot its source resolved to rather than by the source string, so two directories holding the same bytes are the same variant. It is the corpus half of a variant, which also fixes model and effort.

  • Delivery stage — a stage whose artifact is committed code; its evidence is a diff, changed paths, check integrity, and local check results. Today, build.

  • End-to-end mode — running the whole pipeline to evaluate its final committed output. Stage gates still judge intermediate artifacts and can stop execution before a final result exists.

  • Exit code — what the executable returns, with one meaning each: 0 the command completed and wrote its record, whatever the grade; 2 a usage error (unknown flag, missing required flag, unparseable value); 3 a refused precondition (a needed approval whose flag is absent while stdin is not a TTY, a run that cannot be replayed); 1 an execution failure. A failing grade is evidence, not an error.

  • Grading definition digest — the identity of the check list and state scorer that produced an assessment: a SHA-256 over the canonical form of {checks, stateCheck}. Lineage cannot serve, because it hashes the transcript, fixture, prompt, tools, settings, agents, project files and state scorer but not checks, so two cases differing only in a reply check share a lineage. Two assessments of one attempt carrying different digests is how an operator sees that the definition changed between them.

  • Fired reply — the end-of-turn reply João answered with /brief. It is the reply the style produced and he rejected, not the one he wanted; the rewrite he accepted comes later in the same session. A cut is the fired reply's own index, so the prefix keeps everything that produced it and drops the reply itself, and the attempt writes its own reply in that place.

  • Fixture history — the committed git history a session case's fixture carries, seeded into the attempt directory alongside the fixture tree so every arm starts from the same commits. It is stored as a dot-git directory, because git refuses to commit a nested .git, and the seeding renames it and recreates the empty refs/heads and refs/tags that the commit dropped. Without those, git declines to read the seeded directory as a repository and searches upward, so the session reads whatever history encloses the attempt directory, or none.

  • Fork — copying a transcript prefix into the attempt directory's project slug under a fresh uuid, with every occurrence of the source session id rewritten, so a session can be resumed from it without its original working directory. On claude 2.1.258 the resumed session keeps that uuid and appends to the forked file rather than writing a new one.

  • Fresh checkpoint chain — a replay's consumed checkpoint chain when none of its checkpoints is stale.

  • Human review — the verdict, summary, and classified findings a reviewer records against a run's Judge result, in <run>.review.json. The reviewer is a person or the agent standing in for one; the name says whose judgment the record carries, not which hand typed it. rehearse review writes it from flags or from a file, and calibration reads it.

  • Interrupted run — a run whose process ended (a kill -9 or a crash) without writing a terminal artifact or stop record. A signal the handler catches is not one: it stops the pending stage, writes that stage's stop record and the run's operator-stop.json, on the way out. No file on disk gives such a run a status of its own, though the stage it died in may be left as a stage awaiting judgment. The server's startup reconciliation pass finds the run by checking whether the pid the run's claimed target recorded is still alive, and if not, marks the run's event stream run-interrupted, distinct from FAILED, which a graceful signal handler still writes on its own.

  • Judge — evaluator attached to a stage transition: deterministic check or rubric-scored LLM with rationale.

  • Judge agreement baseline — accumulated binary Judge and human decisions for one exact Judge model and frozen rubric contract, summarized separately for each rubric criterion.

  • Grade agreement: how many places apart the operator's and the Judge's letters for one stage sit on the scale A, B, C, D, F, which has no E. The design calls a place a letter step. Zero is exact, and "within one step" counts exact matches too. Distinct from the Judge agreement baseline, which counts binary decisions per rubric criterion. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-273).

  • Judge drift: for one rubric dimension under one Judge model and rubric, the mean of the operator's letter place minus the Judge's over every stage the operator graded. Positive means the Judge grades more generously. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-273).

  • Judge attempt — one Judge call against frozen evidence and a rubric, recording its returned payload, call cost, and whether harness validation accepted or rejected it.

  • Judge progress — while a stage Judge call is in flight, each blocker, requirement or dimension it has finished writing and that passes the checks one item can face, counted against the rubric per section ("4 of 4 hard blockers evaluated", "3 of 5 dimensions returned"), with each returned blocker's PASS or FAIL and each returned dimension's grade. A rejected attempt withdraws its count, and the next attempt counts from none. Progress is best-effort, like every run event, and is never a record.

  • Lineage — hash of everything that produced a checkpoint: upstream checkpoint, corpus files feeding the stage, model, effort, and canonical stage settings.

  • Locator — where a quoted span sits in the source the record holds: a line range, a diff hunk, a commit subject, or a transcript exchange's character range. The harness computes it and never takes it from the Judge. Evidence the harness writes, and evidence citing a harness result, carries a locator naming that result instead of a span.

  • Materialize — write a checkpoint's frozen state into a directory, byte-faithfully, so a stage can run from it.

  • Model family — a named Claude model line — Opus, Sonnet, or Haiku — recognized from either its native alias or a full model ID.

  • No reply — the outcome of a session attempt whose envelope carried no result and did not mark itself an error. A budget halt is marked one, so it is a failed attempt rather than No reply. It is not a reply of zero words: no check is evaluated and none is recorded, so the attempt reads as a measurement that did not happen rather than one that passed.

  • Observed delivery — saved transcript evidence that a Read result or Skill companion placed recorded text into session history. An invocation alone is not a delivery, and delivery does not show that the model followed the text.

  • Review pause — the interactive stop a run makes with the candidate still in the target, asking the reviewer to edit files and press Enter until the calibration validates. It is requested by --pause and needs a TTY, refused before any paid work without one. A run without --pause never stops: it writes the preliminary artifact, retains the candidate, restores the target, and exits, leaving the review and the calibration to their own commands.

  • Pause after step: the operator's request that a running pipeline run start no further stage once the current one is judged and its checkpoint written. The run then restores the target and ends as paused, listed as PAUSED:<stage>. A paused run is not resumed. Continuing it is a replay from that checkpoint. Distinct from the review pause, which holds a candidate in the target for a human edit.

  • Operator stop: a run, replay, group or session attempt the operator ended, by Stop & restore repo in the browser or a signal from a terminal. It kills the commands, restores the target, and records the stop so the run reads OPERATOR_STOPPED rather than failed. Distinct from a ceiling stop and from a stop below the minimum grade.

  • Operator grade: the operator's own judgment of one stage the Judge graded, given per rubric criterion, as the Judge grades, before anything the Judge returned is shown: PASS or FAIL for each hard blocker and requirement, a letter for each dimension, and an optional note. Its stage letter is derived by the same rule as the Judge's. It changes no grade, and calibration does not read it. Distinct from a human review, which classifies findings against a run's Judge result. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-273).

  • Output summary: the phrase a stage's node on the live monitor shows once the stage has finished: what the stage's record says it produced, as counts of its commits, changed files and workflow-state changes, derived with no model call. Where the run has a root-cause analysis, the newest one's phrase for the stage replaces it. Until a stage has either, its node reads "contribution pending". A UI concept; nothing stores it.

  • Pipeline — the ordered stages and their judge attachments, declared as data. The UI's design calls this a task (see UI vocabulary); the word in code, records, and this glossary stays pipeline. Graded as a whole, a pipeline's grade is computed from its first input and its last artifact only, never by averaging stage grades, and a pipeline that stopped early is not gradable as a whole.

  • Pipeline definition — the declared, user-authored data the harness reads to know which stages to run, in what order, and with what skill, expected artifact, and rubric.

  • Planning stage — a stage that produces planning artifacts such as acceptance criteria, card updates, or an attached document. The declaration decides which outputs are required. The bundled pipeline's planning stage is shape.

  • Project instructions — the instruction file a repository carries in its own tree for agents working in it (CLAUDE.md or AGENTS.md). A property of the repository, never installed by the harness. Distinct from the corpus's global CLAUDE.md, which is the file under evaluation.

  • Product Owner (PO) — the dynamic agent that answers stage questions from the product brief; one session per run.

  • Provider call — one invocation of the model provider by a worker, Product Owner, or Judge. Its evidence may include usage metrics; the call remains explicit when those metrics are absent.

  • Quoted span — the text a Judge copied from one source it was given, stored on the evidence item beside the claim. A quote its cited source does not hold rejects the Judge attempt.

  • Regrade — one re-evaluation of a saved attempt's evidence against the case as it stands now, reaching no provider. It reads the recorded reply, the transcript beside the attempt, and the preserved state evidence, and writes an assessment. A check whose evidence the attempt does not hold is reported unavailable rather than graded, and the saved reply, transcript and state evidence are left byte-identical, so a corrected check costs no second paid session.

  • Record ID — how a session names one recorded thing to the CLI and how the CLI names it back: a kind prefix and the identity that kind already has on disk, case:<id>, run:<name>, checkpoint:<run>/<stage>, attempt:stage:<lineage>/<timestamp>, attempt:session:<case>/<uuid>, group:<group-id>, rep:stage:<group-id>/<rep-id>/<stage>, rep:session:<group-id>/<rep-id>, comparison:<manifest-digest>. Every id list prints is one show accepts, and the prefix is parsed once at the boundary into the kind, so show never guesses which record a bare string named. show also accepts a short id, resolved through its case's registry to the Record ID it aliases.

  • Record summary — the short markdown a session pastes onto a card, computed as a pure function of one parsed record: for a run its stages, grades, verdict, and cost; for a group its reliability summary and cost, and its reps' reads, whose states are judged against the linked corpus when show runs, so that one section follows the corpus and names it; for a comparison its per-case paired deltas beside the control arm, or, for a single-case comparison, such as a session or a one-checkpoint stage, its sampling unit, each arm's own interval, and the unpaired contrasts between them. It is never a second record shape: --json still prints the strict record's own bytes.

  • Rep — one repetition of a run; scores are distributions over reps, never a single rep. The UI's design's singular attempt already matches this glossary's Attempt entry and needs no mapping; the design's plural attempts inside a group is this glossary's rep (see UI vocabulary).

  • Rep outcome — one binary reliability observation. A stage succeeds with Judge grade A or B, a final judgment succeeds with PASS, and a session succeeds when its declared checks pass with the required evidence. A stop, no reply, execution failure, or required missing metrics makes a confirmation rep unsuccessful. A comparison records this success value with the judged grade, or records EXECUTION_FAILED, METRICS_MISSING, or NOT_REACHED without a grade. Lowering a continuation threshold does not redefine success.

  • Replay — re-running one stage from a checkpoint with the current corpus, in a fresh worktree.

  • Paired rerun — the replay an instruction edit asks for so its effect can be read: the stage that read the edited file, replayed from the checkpoint it started from against the edited corpus, to set beside the result the edit made stale. Offering it starts nothing.

  • Retained candidate — the run's final result commit, pinned in the target repository under refs/rehearse/<run> before the target is restored, so the candidate outlives the run that produced it. Restoring makes the commit unreachable and only the ref keeps gc from pruning it; show run:<name> --checkout <dir> materializes it as a detached worktree. It is the same ref a checkpoint is pinned under, named by the run rather than by a stage.

  • Root-cause analysis: one sealed agent session's reading of which corpus file, if any, an ended pipeline run's outcome traces to, kept as a record of its own beside the run's records. It is an opinion read from recorded evidence, never a measurement, and a run can hold several. An analysis is in flight while its session is still running, distinct from a run in flight. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-421). It gives each declared stage one role:

    • Root cause: the corpus file the analysis names as what the outcome traces to, and the role of the one stage that read it. No stage has this role when the analysis names no file.
    • Contributing factor: a stage the analysis reads as having moved the outcome without being its root cause.
    • Not a factor: a stage that ran and that the analysis reads as not having moved the outcome.
    • Never ran: a declared stage whose agent session never started, because the run stopped before it or the spend ceiling refused it, recorded by the harness rather than read by the agent.
  • Rubric — the frozen grading contract a Judge applies; per-stage under the case's rubrics/, final in the case's rubric.md.

  • Rubric criterion — one identified hard blocker, requirement, or quality dimension within a rubric, reduced to a binary pass/fail decision for calibration.

  • Score — a statistical summary over a confirmation run's rep outcomes: their distribution, success rate with standard error, and pass^k. A single-rep Judge result is evidence, not a score.

  • Partial score — how many of a rep's declared graded outcomes held, which separates a rep that missed one outcome from one that missed them all where a rep outcome alone cannot. A comparison report carries two such tallies per rep, one over the declared checks and one over the declared state results, each counting what passed against what was declared and listing what failed: a failing check by its declaration index and kind, a failing state result by its declared name. No field combines the two, show prints neither, and neither changes the rep outcome, which the live path takes from the checks alone.

  • Run artifact — the recorded evidence of a run under .benchmark-runs/.

  • Run artifact transition — one persistence operation that advances a run's main or stage record. Transitions are serialized; abort recording is terminal and cannot be overwritten by a later normal transition.

  • Run event — a timestamped progress fact appended to .benchmark-runs/run-events.sqlite, keyed by run ID and streamed through the server's SSE endpoint. It is best-effort derived progress state. JSON artifacts remain authoritative for recorded conclusions; there is no event-history rebuild command.

  • Run in flight — a run that is executing right now: its event stream's latest entry is non-terminal, no artifact or stop record has been written for it, and the process that claimed its target is still alive. The run-history report gives such a run the status RUNNING. The liveness check is what separates it from an interrupted run, whose stream also ends non-terminal. A target holds one claim at a time and that claim names no run, so the check answers for the target: a crashed run whose target a later run has claimed can still read as in flight.

  • Run spend — what a whole run has paid, across worker sessions, Judges, and the Product Owner. Every run event carries what the run has paid so far, except events recorded before run events carried it, and a terminal run event, run-completed or a run-failed with a persisted artifact, carries the whole. Distinct from stage spend and from the per-session limit the session knobs set.

  • Burn rate: a run in flight's run spend divided by its elapsed time, as the latest run event measured them, read in dollars per minute. A UI concept only: the live monitor computes it and the harness stores none (see screen 2b of the design spec). Name accepted unattended as unsettled, pending the operator's confirmation (ACT-270).

  • Remaining estimate: the time and spend a run in flight is expected to take before its steps finish. The time sums each unfinished step's median graded time in earlier runs of the same case, the running step's less what it has already run and never below zero, and the spend is that time at the burn rate. When a figure it needs is missing there is no estimate, and the reason is shown instead, from the list under Stage times in the reference. A UI concept only: the harness serves the median times and the live monitor computes the estimate. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-270).

  • Spend ceiling: the stored USD limit on a run's whole spend, which every paid command requires before it starts. Each session a run starts gets a budget no larger than the ceiling minus the run spend so far, and once the spend reaches it the run stops in the stage where that happened. It can be overrun by the calls in flight when it is reached, one per running session. A confirmation group is held to its attempts times the ceiling. Distinct from the per-session limit the session knobs set.

  • Ceiling stop: the stop a run makes when its spend reaches the spend ceiling. A stage stopped this way writes its stop record with the ceiling and the spend, and a final Judge stopped this way fails the run with both on its failed run record. The target is restored and no later session starts.

  • Group ceiling: the spend a confirmation group may reach, its attempts times the spend ceiling. The group stops starting sessions once its spend reaches it.

  • Budget halt: the provider ending a session because its spend reached the session budget passed to it. What it reports spent can exceed that budget. The envelope marks it an error, writes no result, reports the spend and, in the wording every observed CLI version used, states the cap in its errors, so a session attempt ending this way is recorded as failed rather than as No reply. Distinct from a ceiling stop, which the harness makes, though a session whose budget was the ceiling's remainder can halt this way just before the run stops. Name accepted unattended as unsettled, pending the operator's confirmation (ACT-230).

  • Stage spend — what one stage has cost. Every non-terminal run event carries a spend figure, and the event's kind decides which stage spend it is: stage-started reports the stages finished before this one, turn-completed this stage's session so far, stage-judging and judge-progress this stage's finished session, and stage-completed that session together with its Judge. None of them is run spend, which is why a reading of one is shown with the words for what it covers.

  • Unreadable record — a saved record (a pipeline run, a session attempt, a stage replay or a confirmation run) whose own files the run-history report could not read into a row, carried by its kind, its record ID and a reason with absolute paths redacted. A malformed artifact, a run directory with neither a manifest nor run events, or a group directory with no group record produces one. A run that failed or was interrupted before writing a manifest is still listed as a row from its events. An unreadable record is data the report carries, not a failure of the report, so one bad record does not blank the rest. Staleness is judged for every run before any of them is read, so a corpus the staleness pass cannot resolve does not produce an unreadable record: it produces a staleness cause, or no report at all. A staleness report carries its own unreadable list, keyed by case and attempt rather than by run; the two are the same idea applied to different records.

  • Project slug — the name the provider gives the directory it writes a session file into: the working directory's real path with every / replaced by -. On macOS /tmp/x resolves through its real path first, so it is -private-tmp-x.

  • Sealed session — a Claude session with safe mode and no tools, used for judges.

  • Session case — a benchmark case whose unit of work is one Claude session. It declares a prompt, tools, corpus files, checks, and optional fixture, transcript prefix, settings, agents, project files, and state check. It runs once for debugging or as isolated confirmation reps, and either way every corpus file the case declares, skills and CLAUDE.md included, is written under the attempt's .claude/.

  • Session naming — the uuid an attempt gives its own session before the call, as the fork's id when resuming and through --session-id otherwise. It is what lets the attempt name the one session file it owns under its slug, so cleanup deletes that file and never an entry it cannot account for, whether the call returned or threw. Seen from the other side, a transcript's records carry the id of the session that wrote them, whatever the file or directory holding it is called, and that is the identity a capture reads and a fork rewrites. A preserved attempt's transcript therefore names the session the provider ran, not the uuid of the directory it was saved under. A pipeline stage names its session the same way, and its start records the name, so the live monitor can read the transcript while the session writes it.

  • Session knobs — the CLI and environment settings shared by run and replay that select the workflow and Judge models and efforts and set the per-session budget.

  • Session tail: the last lines of a running stage's transcript as the live monitor's session pane shows them, each a user or assistant message, a tool call collapsed to one line, or a tool result. A UI concept only (see screen 2d of the design spec).

  • Short id: how the operator and a session name a run, replay, session attempt, confirmation run or checkpoint, scoped by its case: <case>/r<n>, <case>/g<n>, <case>/r<n>/s<k>. Each case numbers its runs, replays, session attempts and confirmation runs in one sequence, and a number once given names nothing else. s0 is the checkpoint taken after task setup and s<k> the one after the k-th stage of the run's frozen pipeline. It is an alias for the record's Record ID, which stays canonical. A replay or confirmation rep also has a place among the attempts at its checkpoint, shown as attempt n of m, where the original run's stage counts when it recorded a result there. The place is a label and not a name, since m grows with every later attempt and n can too. A rep of a session-mode group has no checkpoint and is placed by its position in its group.

  • Stage — one pipeline step: a skill invocation consuming upstream artifacts and emitting its own. The UI's design calls this a step (see UI vocabulary); the word in code, records, and this glossary stays stage.

  • Stage awaiting judgment — the stage an interrupted run was in after its session finished and before its judging completed. Its <run>.<stage>.json carries AWAITING_STAGE_JUDGE, the stage name, the Judge's input and the model, effort and budget the stage ran with. Judging ending overwrites that file either way, with a scorecard or with a stop record, so a file still carrying the status is one whose run died in that window. It writes no checkpoint, so it has neither a lineage nor a stop reason, it declares no corpus files, and the parsed exchanges its input holds are not the raw transcript. It is not a stopped stage: nothing judged it and nothing failed.

  • Stage commit history — oldest-first subjects of the commits a stage added after its baseline; absent when the stage did not advance the target history.

  • Stage corpus reconciliation — the per-declared-file comparison between a checkpoint's declared corpus files and the reads its stage transcript shows. Each entry reads observed, no observation recorded, or undeclared. It is keyed on the corpus layout path and never on the declared hash, since a hash of bytes on disk cannot establish that their contents entered context. No observation recorded means no recognized read, not absence from context.

  • Stage scorecard — persisted Judge result for one stage: its frozen input and rubric, citations, grade, prompt, and Judge cost; a rejected scorecard also carries its calibration.

  • Stage settings — the harness-owned JSON a stage session is started with, holding a permissions deny list and a fixed set of boolean feature switches. It is never a copy of the operator's live settings, and the schema refuses every other key, a hooks block included. A pipeline case may name its own file, and a case that names none gets the committed root stage-settings.json. A checkpoint records the file's canonical digest, the hash of the re-serialized JSON the session was given rather than of the file's raw bytes. Staleness compares that digest alone, while lineage folds in the recorded path beside it. See the reference for what a record stores and when an edit stales a run.

  • Stale checkpoint — a checkpoint whose recorded inputs (corpus files, stage settings, model, effort, its judge rubric, or an upstream checkpoint) no longer match the current state; still replayable for exploration, refused in comparisons.

  • Stale provenance — the state of a comparison rep whose saved evidence the comparison cannot vouch for: its confirmation run, rep or attempt file is missing, unreadable, differs from the digest the comparison recorded, sits somewhere other than where its confirmation run keeps it, or names a different case, run or rep. The comparison page still shows the rep, marked stale, but offers no path into its attempt history. A rep can be stale from the moment the comparison is saved, when its files were never at their confirmation run's own location. It is unrelated to a stale session attempt or a stale checkpoint, which compare a record against the current inputs rather than against the files it was measured from.

  • Stale session attempt: a session attempt whose recorded corpus file digests the corpus under test no longer matches. It is the session kind's counterpart to a stale checkpoint, the same claim that a measurement no longer describes the corpus, keyed on the files the case declared rather than on the skill a stage invoked. Each attempt is judged, not only a case's most recent one.

  • State check — the grading definition a session case declares for the files and git state its session leaves: a command to run and the outcome names it must report. It is declared inline in case.json, which is what folds it into the attempt's lineage, and it runs against a restored copy of the attempt state evidence rather than against a live tree. It reports one result per declared outcome, keyed by the name the case declared where a reply check's result is keyed by kind, and those results are recorded in their own field rather than in the check list, whose members are all evaluated against the reply and the transcript. Before the command runs, the case's own copy of every path it names is laid back over the restore, so a session that rewrote the scorer is still graded by the case's bytes.

  • State grading error — the record a session attempt carries when its declared state check could not produce grades: the scorer would not run, exited non-zero, printed output the result schema rejects, or omitted a declared outcome. It is a distinct fact from a failed grade, because none of those says anything about the session's work.

  • Staleness cause — one named statement of why a recorded result no longer describes the current state. A cause names the thing that moved, an upstream stage, the model, the effort, the stage settings file, one corpus file that changed, was added, or was removed, or the judge rubric, or else it carries a corpus refusal verbatim, since a corpus that cannot be read cannot be shown to still match. A checkpoint carries its causes as a list: those that are not about corpus files, its judge rubric among them, then one per corpus file that drifted, so the list has no bound but the corpus's size. A replay lists its judge rubric last. Its length counts reasons and is not a distance between corpus versions.

  • Stage kind — which validation and evidence strategy a stage uses: planning or delivery. Declared per stage, independent of the stage's name.

  • Stage mode — running one stage against frozen upstream artifacts. This supports local debugging and repeated measurements, but an intermediate grade remains a proxy for the quality of the complete workflow.

  • Stop record — the <run>.<stage>.json a stopping stage writes in place of its scorecard, carrying the stopping status, the stage name, the reason the run ended there, the input the Judge was given, and the model, effort and declared corpus files the stage ran with. A stop at a grade below the minimum grade also carries the letter and verdict, the minimum grade, the Judge's attempts, the stage and run elapsed times, and the Product Owner's cost and calls up to the stop; older stop records lack them. The run's manifest still supplies the case; only the checkpoint the stage never wrote is missing.

  • Stopped stage — the stage a run ended on, whether its grade did not meet the pipeline's minimum, its judging failed, or a signal stopped the run mid-stage. Only the reason its stop record carries says which of those it was. It writes no checkpoint, so it has no lineage placing it among the run's other stages; its stop record names it by run and stage instead. The run list shows the run as STOPPED:<stage>, linked to that stage's context history.

  • Target repository (template project) — the real application repository, kept at a stable baseline, that tasks run against.

  • Target check — one command declared by the pipeline and run against the target repository both at baseline and after delivery.

  • Task graph — the UI's horizontal chain of stage-node cards (grade, status, live tool call, checkpoint, output summary, in/out counts of instruction files loaded and artifacts produced) shown on the live monitor and, in reduced form, as "the map" on the run-detail Contribution layout. A UI concept only; nothing in the harness computes or stores a graph (see UI vocabulary).

  • Trajectory step — one workflow-agent turn reported by the provider. PO and Judge turns are excluded so the measure tracks corpus-induced workflow behavior.

  • Transcript prefix — a real session file truncated at a cut and used as frozen starting context. It lives in the case directory the declaration belongs to, is checked against the declared digest before an attempt resumes from it, and is carried in lineage. A prefix holding private material is kept out of the repository, so the case reaches another clone without its bytes.

  • Transcript diagnostics — the compact projection saved on each new session attempt record from transcript records at and after its captured cut. It records source and measured line counts; raw tool-use occurrences; explicit is_error: true tool results; and exact repeated Bash.input.command values on distinct unique tool-use IDs. Locations use 1-based JSONL lines and content blocks in the whole retained transcript. Repeated commands keep a digest, character count, bounded preview, truncation flag, and ordered locations; the complete input stays in the raw transcript. complete, partial, and unavailable distinguish an observed zero from incomplete or missing evidence. Repetition does not mean waste and carries no phase, token, cost, or causal attribution. An absent field on a historical record means the projection was not recorded; readers do not reconstruct it from today's case declaration.

  • Total input tokens — one model request's reported input_tokens + cache_read_input_tokens + cache_creation_input_tokens. The three categories are disjoint in every saved transcript, so the sum double-counts nothing. The provider's prompt-caching documentation names this sum total_input_tokens; no saved record carries the sum as a field, so the harness computes it. It excludes the request's own output tokens, which rejoin the prompt on the next request. It is not the active context window: no saved record carries an occupancy figure to render it against, and the window limit it would be rendered against comes from the per-model usage block rather than from the transcript. The series across an attempt does not rise monotonically: a cache re-warm moves tokens from cache_read_input_tokens to cache_creation_input_tokens and lowers the sum, observed in 6 of 54 saved attempts, and rows recording no model request report every category zero.

  • Attempt cost reconciliation — three separately stated readings of one attempt's cost, taken from its saved transcript and record rather than from a supplied context evidence bundle, whose own per-request pricing is described in reference: the cost the provider reported on the attempt record, the sum of per-request calculated costs over attempt-region requests only, and the difference between them. They are never folded into one figure, because the gap between a provider charge and a catalog-derived sum is evidence about the catalog. Each reading is complete, incomplete, or unavailable with its reasons, and a reader takes the state before the figure: an incomplete reading still carries a number, and that number is a partial sum rather than a total. A request whose usage, model, cache-write TTL split or rate is missing or in conflict leaves the calculated reading incomplete and names why, rather than passing as a priced request costing zero.

  • Per-model usage block — the provider's own account of one CLI call, broken down by the models the call used. It carries each model's input, output, cache-read and cache-creation tokens, the cost the provider charged for them, the model's context window and maximum output tokens, and the basis that cost was priced on. The harness retains it verbatim on an attempt's metrics. A call made by a CLI that reports no such block records its absence, which is not the same as a call that used no model.

  • Served model: the model that actually answered a call, as distinct from the model the call requested. The two can differ, for example after a provider's safety-classifier retry. The per-model usage block names it; Rehearse records that block but does not read the served model from it.

  • Rate catalog — the per-model, per-category prices a request's cost is calculated from, carrying the source and version they came from. A saved calculation persists the catalog that priced it, so a later catalog changes what new calculations cost and leaves the saved ones alone. A reading computed on demand persists nothing and prices from whatever catalog its caller supplies, reporting itself unavailable when none does. Each category records where its price came from, because the suite can only defend a rate the saved attempts re-derive. A category no saved attempt exercises is marked for what it does rest on, a price someone looked up or a figure computed from another category, and pricing that rests on either is weaker evidence than pricing the corpus measures.

  • Instruction load — one instruction file an attempt loaded automatically, named by its path and the kind of memory it came from. Two sources record loads and they carry different detail. A hook capture also records why the file loaded, what triggered it and which file included it. A transcript records none of those three, so a load read from a transcript reports them unavailable rather than guessing. An attempt whose source records no loads at all is distinct from one that loaded none.

  • Cost basis — what the provider priced a model's reported cost on. A cost basis of list is public per-token rates, so the cost can be re-derived from a rate catalog and checked. Any other basis is a discount or plan the catalog does not describe: the cost remains the spend the provider reported, and it is not evidence about rates.

  • Variant — a named configuration: corpus snapshot, model, and effort.

  • Version distance: how many corpus versions a record sits behind the corpus under test: 0 when the corpus under test is the version the record measured, and otherwise the positions between that version's latest entry in the corpus's version log and the corpus under test. It is not recorded for the initial checkpoint, which reads no corpus, a record written before versions, one whose version is absent from the log, or a corpus that refuses an entry. It never makes a record stale on its own. Only a staleness cause does. Name accepted unattended as unsettled, pending the operator's confirmation (doc-157, question 6).

  • Version log: the ordered corpus versions one corpus source has been measured at, in one records directory. A version enters only when it differs from the source's latest entry, and order is position, not time. Name accepted unattended as unsettled, pending the operator's confirmation (doc-157, question 6).

  • Workflow state — the backlog/ and .boris/ trees copied independently of Git to carry workflow artifacts across stage materialization and target restoration. A target's Backlog configuration determines where its board lives. These target artifacts are distinct from Rehearse's external personal board.

  • Attempt elapsed time — the wall-clock duration of one attempt, from its start to its finish, including work outside provider calls. Recorded per rep, and carried into a comparison as each arm's observations and their mean. The observations are sorted by duration rather than by rep ordinal, so they describe the arm's spread and not which rep was slow. It is not the sum of a call's provider durations, and summing it across attempts that ran concurrently does not give wall-clock time.

  • Provider duration — the time spent inside provider calls. The attempt elapsed time of the same attempt encloses it, because it also counts the work around those calls. Comparison reports do not carry it.

  • Group makespan — the wall-clock duration of a confirmation group, from its first attempt starting to its last finishing. Smaller than the summed attempt elapsed time whenever attempts ran concurrently. Recorded on the group record and not carried into comparison reports, because it describes how the operator scheduled the reps rather than the treatment under test.

  • Group median grade: the grade a confirmation group's reps received at one stage, taken from the reps a Judge graded there. On an even count it is the lower of the two middle grades, so it is always a grade some rep received. A session group has none, because its reps pass or fail on checks and no Judge gives them a letter, though its row still counts them. Rule accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 3).

  • Case median verdict: the median of the final Judge's verdicts over a pipeline case's runs at one corpus version, FAIL ranked below PASS. On an even count it is the lower of the two middle verdicts, as a group median grade is, so it is always a verdict some run received. A session case has none, and its figures count the runs that passed their checks instead. Rule accepted unattended as unsettled, pending the operator's confirmation (doc-198, decision 5).

  • Stage elapsed time — the wall-clock duration of one stage in a pipeline run, from the stage starting to its Judge's grade, read from the run's clock. Recorded on the stage's scorecard, and on its stop record when the stage's grade fell below the minimum grade. Name accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 4).

  • Run elapsed time — the wall-clock duration of a pipeline run, from its start to its main artifact, or to the stop for a run a stage's grade stopped. It encloses every stage elapsed time. Name accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 4).

  • Silence limit: the longest a stage's provider call may write nothing to its stdout before the harness ends it, 30 minutes. A call that keeps writing runs as long as it needs. Name and length accepted unattended as unsettled, pending the operator's confirmation (ACT-347).

  • Minimum grade — the letter every stage of a run must reach for the run to continue, set by --minimum-grade and B by default. The run manifest records it. Changing it never changes the grade a Judge recorded. Name accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 4).

  • Run cost — the sum of the costs a pipeline run's records hold: each stage's session and Judge, the Product Owner, and the final Judge, naming each part it summed and each it lacks. Distinct from run spend, the figure a terminal run event carries. Name accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 4).

  • Final outcome — what a pipeline run's final Judge recorded, the UI design's task grade: judged PASS or FAIL, judging failed, pending while the run is live, or not reached with the reason the run ended. The final Judge returns no letter. Name accepted unattended as unsettled, pending the operator's confirmation (doc-152, decision 4).