Find out whether changing your coding agent's instructions actually helps.
Rehearse runs a fixed task through the real Claude Code CLI, records the instructions and evidence, and grades the result. You can replay a workflow stage after editing a skill, or repeat a session to see how much its results vary. The goal is a debugger and regression suite for an engineer's instruction corpus: project guidance, skills, output styles, and agent definitions.
For example, a shorter instruction might produce a better reply once. Rehearse helps you inspect that attempt, repeat the task, and compare quality with cost before deciding the instruction earned its place.
Early development. The CLI is usable, and a local browser UI reads recorded evidence. Some workflows still depend on the maintainer's environment. Session confirmation groups can feed comparisons alongside stage and pipeline groups, over several cases or over repeated runs of a single one. See current state and priorities before planning an experiment.
Choose a case + instructions + model
|
Run the real agent
|
Grade and retain evidence
|
Inspect → edit → run again
|
Confirm with repeated trials
- Session cases run one prompt in a temporary directory, optionally with fixture files or a captured conversation prefix. Deterministic checks grade the reply and tool transcript.
- Pipeline cases run a declared sequence of skills against a target Git repository. A Product Owner answers questions, independent Judges grade each stage, and checkpoints let you replay a stage without rerunning its predecessors.
- Confirmation runs repeat frozen inputs and report reliability and resource usage. A single attempt is debugging evidence, not proof of improvement.
The current provider is Claude Code. This is a local tool using your credentials
and agent environment, with records under .benchmark-runs/.
Install mise, then from a clone:
git clone https://github.com/joaofnds/rehearse.git
cd rehearse
mise install
mise exec -- bun install --frozen-lockfile
mise exec -- bun run rehearse --help
mise exec -- bun run rehearse case list
mise exec -- bun run rehearse case show smoke --jsonThese commands inspect the project without calling a model. The pinned toolchain is in mise.toml. Install and authenticate Claude Code separately using its quickstart; it is needed only when you run an experiment.
Experiments whose sessions can run commands, such as every pipeline run and
replay and a session case that declares a tool or hooks, run only on macOS.
Rehearse keeps those sessions from directly signalling your other processes with
/usr/bin/sandbox-exec, and refuses on a host without it. The
reference says what that does and does
not cover.
To run your first experiment: follow the runbook. It
provides the corpus file the smoke case needs, explains the session budget, and
shows how to read the result. A bare run selects the audit-log pipeline, so
choose --case smoke explicitly when starting out.
To browse the UI: build and serve the client, then open
http://localhost:4173.
mise exec -- bun run build:client
mise exec -- bun run serveA fresh clone has no run history. Available views cover run history, the linked
corpus, saved comparisons, saved session-attempt history, the tasks and cases
the checkout declares, the settings, and the design system. Declare a case on the Cases view
writes an uncommitted cases/<id>/case.json into the checkout.
The run history shows a run that is executing, with its stage, elapsed time, and
spend, and the Live monitor follows it live. New run and Replay a step on run history and Replay
from here on a pipeline stage start a run or replay under the stored spend
ceiling, which spends provider money. Several prototype screens are still
planned.
The server is intended for local use, binds to IPv4 loopback, and has no
authentication.
| I want to… | Read |
|---|---|
| Understand the goals and long-term direction | Vision |
| Understand context visibility and efficiency plans | Context assessment and roadmap |
| Know what works and what needs work | Current state and priorities |
| Run an experiment and inspect its evidence | Runbook |
| Configure cases, replay, grading, and comparisons | Harness reference |
| Understand the implementation | Architecture |
| Contribute code or keep the docs current | Contributing |
| Look up a project term | Glossary |
| Read the evaluation methodology | Research |
The documentation index also explains the status of the UI design and recovered historical material. Public project context lives in these tracked documents; the maintainer's personal task board is outside this repository.
Licensed under the Apache License 2.0.