prosaic's harness has three layers, all run by pytest:
uv run pytest # everything (AI checks included if an agent CLI is present)
PROSAIC_AI_TESTS=0 uv run pytest # deterministic only
uv run pytest -m ai # only the AI-judged checks
uv run pytest -q <path> # a narrow selectionCI runs four commands, in this order, and a green uv run pytest does
not imply the first three pass:
uv sync --locked # what CI installs from -- see the warning below
uv run ruff check .
uv run ruff format --check .
uv run mypy tests
uv run pytestuv run does not prune the environment. It installs what is
missing; it does not remove what is no longer declared. So after a
dependency is removed, a machine that already had it keeps passing
while CI -- which runs uv sync --locked -- fails. This is not
hypothetical: removing the typed library took pydantic out of the
dependencies and left plugins = ["pydantic.mypy"] in [tool.mypy],
which passed locally against a stale venv and aborted mypy in CI before
it checked a single file. Two more stale-config failures were hiding
behind the same venv and surfaced the moment it was pruned.
Run uv sync --locked before trusting any dependency removal. The
statically detectable half of that failure class is now pinned by
tests/test_ci_config.py -- config naming a path, plugin, or flag that
no longer exists -- but the environment itself is not something a test
can check from inside a process running in it.
Note also that mypy's exclude stops files being collected, not
followed: a type-checked file that imports an excluded module pulls
that module into the checked set. Adding one test that imported
tests/harness/ai.py is what first type-checked it.
Fast, deterministic: fit math, caption parsing, descriptor↔blank consistency for every registered form (the revision-drift alarm), smoke fills, overflow plumbing.
The heart of the harness. A scenario is an entire fictional matter
checked in as a fixture (matter/ — sources, config, front matter) in
a fixed starting state. The test copies it to a temp dir, performs
real operations through the system (engine fills, envelope builds;
syncs and triage as those grow fixtures), and then makes many
independent checks on the results. Each scenario names the spec it
executes; when a spec promise gains a test, the spec's (untested)
marker comes off.
Current scenarios:
form_filling— a motion matter; checks field placement, caption repetition, mandatory blanks, overflow → MC-025, checkbox linkage, cover-sheet assembly, registry-wide invariants; plus AI visual QA.pleading_build— a declaration matter; checks pleading-paper anatomy (28 line numbers, caption, perjury clause), typography conventions (em/en dashes, no spaced dashes), annotation leakage; plus an AI filing-readiness judgment.
For properties that are judgments rather than mechanics ("this
rendered form is court-ready"), tests call the judge in
tests/harness/ai.py: a headless agent invocation (through
cli/agent-run, so any supported agent CLI works) that inspects
artifacts (rendered page images, extracted text) against a rubric
and hard-failure conditions, returning
{score: 0–10, hard_failures, rationale}. A check passes at
score >= threshold with no hard failure. Verdicts are cached under
tests/.ai_cache/, keyed on the question asked (task, rubric, hard
failures, threshold) plus each artifact's name and contents — never on
its path, since scenario artifacts are rebuilt into a fresh temp
directory every run. Delete the directory to re-judge; failures print
the judge's rationale so a disagreement is arguable, not mystical. See
design/adr/0008 for why this is in-house rather than an eval framework.
An unreachable judge is not a verdict. If the CLI is installed but
failing — expired login, rate limit, timeout, a reply that is not JSON
— the judgment comes back unavailable and assert_judgment skips
rather than failing. This matters more than it sounds: a scored zero is
a specific accusation ("the redaction leaked"), and reporting that
about work product nobody judged is worse than reporting nothing.
Transient failures are retried with backoff first, and unavailable
verdicts are never cached, so an outage cannot freeze itself into every
later run.
If you read Judgment.passed directly instead of calling
assert_judgment, call skip_if_unavailable(j) first — otherwise an
unreachable judge answers your question for you. The calibration test
is the cautionary case: it asserts the judge rejects sabotaged
output, so an unreachable judge reporting "did not pass" would satisfy
it for exactly the wrong reason.
Writing a good AI check:
- The rubric describes a 10/10 concretely and says what costs points. Vague rubrics produce vague scores.
- Hard failures are for properties that must never ship regardless of overall quality (a pre-filled signature line, clipped text). Deterministic checks should also cover them where mechanically possible — the judge is defense in depth, not the only defense.
- Judge rendered images for visual properties, extracted text for
textual ones; pre-render PDFs with
scenario.rasterize.
Suites that test the tree rather than the product. They exist because each one covers a failure that is silent — nothing breaks, nothing warns, and the gap looks identical to correctness from the inside.
test_repo_hygiene.py— no absolute paths into a user's home directory, no credential shapes. A leak and a portability bug are the same defect here..githooks/pre-pushruns it at the moment a leak would stop being local.test_docs_coverage.py— everyscsubcommand has a promise inspecs/cli.md, no promise outlives its command, ADRs are uniquely numbered and indexed, everydocs/page is linked from somewhere.test_system_dependencies.py— every program the code invokes by name is declared insystem-dependencies.yaml, and every entry still has a call site. Checked in both directions.test_platform_seams.py— directory and credential policy stay behind their single owners (ADRs 0011, 0012).
Two of these walk git ls-files, so they run in a clone and not in
the container image, where .git is deliberately absent.
- Read the spec you're executing (or write it first — specs/README.md).
tests/scenarios/<name>/matter/— build the smallest fictional matter that makes the operations real. Fictional data only (Jane Roe / John Smith / 24CV00000); mark sourcesnotreal:.test_<name>.py— operations viatests/harness/scenario.pyhelpers, then independent checks: deterministic asserts first, AI judgments for the properties only judgment can see.- Reference the spec in the module docstring, and update the spec's (tested) markers.