The problem. Five engineers and CI all run the same behavioural suite. Every one of them pays separately for identical answers, and a full run is slow enough that people stop doing it before merging.
The shape. A content-addressed cache with a shared tier behind the local disk. The first person to ask a question pays; everyone else replays.
A key is the SHA-256 of the request and nothing else — one rule, documented once. That is the whole reason a shared cache is possible.
--8<-- "examples/21-caching-basics/domarinn.yaml"backend |
Shared tier | Use when |
|---|---|---|
disk |
none | Solo work. The default. |
layered |
S3 when cache.s3 is set, else the results server's /api/v1/cache |
Anything shared. |
Those are the two. http and s3 still parse as deprecated aliases for layered and warn at startup, but they name one tier outright instead of letting cache.s3 choose — so backend: s3 with no cache.s3 block degrades to local disk alone rather than falling back to the server. See caching.md.
A remote always keeps the local tier in front, so a warm local hit never touches the network.
/// tip | The config names only the kind
cache:
backend: layeredNo URL, no credentials — those come from the environment (DOMARINN_SERVER_URL, DOMARINN_TOKEN, or the AWS credential chain). That is what makes a suite safe to commit, and it means the same file works locally and in CI with no branching.
If the credentials are missing, domarinn falls back to local disk with a warning rather than failing the run. Convenient, and worth knowing about: a misconfigured CI job will look like it is working while paying full price. Check for the warning.
Uploading a run is the other way round. The cache degrades because a cold cache still answers the question, only slower and dearer. domarinn run --share has no such fallback — the results have exactly one destination — so a failed upload fails the run with exit 3, and a server whose result-schema window excludes the CLI — or no server URL configured at all — is refused with exit 2 before the suite spends anything. Pass --allow-share-failure (the action's allow-share-failure input) where publishing is genuinely optional. See Uploading CI runs.
///
For sharing to hold, keep whatever changes the request identical everywhere: a different model, endpoint, params, header value or cache_salt is a different request, and therefore a different entry nobody else hits. Cosmetic differences do not count — a base_url with and without a trailing slash names one endpoint and keys one way.
Nothing about the environment has to match. Different checkout paths, different file timestamps, a binary rebuilt from the same commit, a different working directory, unrelated exported variables, two teammates holding different API keys — none of these move a key, by construction. That is what makes a shared backend worth having, and crates/domarinn-core/tests/cache_portability.rs pins each one so it stays true.
This is the part that decides whether a shared cache stays useful or gets thrown away weekly.
--8<-- "examples/22-cache-salts/domarinn.yaml"Two levels, two different jobs:
- Provider-level
cache_salt— a coarse "same build?" pin. Bump it when the program's own logic changes. A commit SHA or a release tag. - Per-case
cache_salt: "$digest: …"— a content digest of just what this case exercises.
Do not make the provider-level salt a content digest of everything your program reads. It works, and it discards the entire cache on any edit — which is exactly the outcome the per-case salt exists to prevent. With both in place, editing one prompt re-runs the handful of cases that use it and replays the rest.
The theory — why a salt is a version pin rather than an entry ticket, and how $digest: resolves — is in caching.md.
- One entry per key, immutable. First write wins, on every backend — so concurrent writers are race-free by construction.
- Errors are never cached. Only successful responses.
- Grader calls are cached too, as requests like any other: the grader's HTTP call, an embedding, an
execgrader's round-trip. A warm run re-pays neither the provider nor the grader.--no-grader-cachere-grades while still replaying provider responses, which is what you want while iterating on a rubric. - A
thresholdis not in the key. It is applied on read, so editing a threshold re-scores instantly instead of re-paying the grader. - Pricing is not in the key either.
cost_usdis recomputed on every hit, so correcting a rate re-prices history rather than discarding it. latencyassertions bypass the cache entirely — a replayed response has no honest latency — and under--cache-onlythe case is refused rather than called live.
The full list is in caching.md.
$ domarinn run eval/behavioral.yaml
$ domarinn run eval/behavioral.yaml # second runThe second run should report every cell as a cache hit. If it does not, something in the suite varies between runs — a now() in a var, a $digest: glob matching a file that is being rewritten, or a salt containing a timestamp.
$ domarinn cache stats eval/
$ domarinn cache gc --older-than 30d eval/- Caching — the full key semantics and backend details.
- Example 21 and 22.
- Server — running the shared tier.