Skip to content
joaofndsPublic

About

Benchmark and improve coding-agent instructions with real Claude Code sessions, stage replay, and graded evidence.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Rehearse

Find out whether changing your coding agent's instructions actually helps.

Rehearse runs a fixed task through the real Claude Code CLI, records the instructions and evidence, and grades the result. You can replay a workflow stage after editing a skill, or repeat a session to see how much its results vary. The goal is a debugger and regression suite for an engineer's instruction corpus: project guidance, skills, output styles, and agent definitions.

For example, a shorter instruction might produce a better reply once. Rehearse helps you inspect that attempt, repeat the task, and compare quality with cost before deciding the instruction earned its place.

Early development. The CLI is usable, and a local browser UI reads recorded evidence. Some workflows still depend on the maintainer's environment. Session confirmation groups can feed comparisons alongside stage and pipeline groups, over several cases or over repeated runs of a single one. See current state and priorities before planning an experiment.

How it works

Choose a case + instructions + model
                  |
          Run the real agent
                  |
       Grade and retain evidence
                  |
       Inspect → edit → run again
                  |
     Confirm with repeated trials
  • Session cases run one prompt in a temporary directory, optionally with fixture files or a captured conversation prefix. Deterministic checks grade the reply and tool transcript.
  • Pipeline cases run a declared sequence of skills against a target Git repository. A Product Owner answers questions, independent Judges grade each stage, and checkpoints let you replay a stage without rerunning its predecessors.
  • Confirmation runs repeat frozen inputs and report reliability and resource usage. A single attempt is debugging evidence, not proof of improvement.

The current provider is Claude Code. This is a local tool using your credentials and agent environment, with records under .benchmark-runs/.

Start here

Install mise, then from a clone:

git clone https://github.com/joaofnds/rehearse.git
cd rehearse
mise install
mise exec -- bun install --frozen-lockfile
mise exec -- bun run rehearse --help
mise exec -- bun run rehearse case list
mise exec -- bun run rehearse case show smoke --json

These commands inspect the project without calling a model. The pinned toolchain is in mise.toml. Install and authenticate Claude Code separately using its quickstart; it is needed only when you run an experiment.

Experiments whose sessions can run commands, such as every pipeline run and replay and a session case that declares a tool or hooks, run only on macOS. Rehearse keeps those sessions from directly signalling your other processes with /usr/bin/sandbox-exec, and refuses on a host without it. The reference says what that does and does not cover.

To run your first experiment: follow the runbook. It provides the corpus file the smoke case needs, explains the session budget, and shows how to read the result. A bare run selects the audit-log pipeline, so choose --case smoke explicitly when starting out.

To browse the UI: build and serve the client, then open http://localhost:4173.

mise exec -- bun run build:client
mise exec -- bun run serve

A fresh clone has no run history. Available views cover run history, the linked corpus, saved comparisons, saved session-attempt history, the tasks and cases the checkout declares, the settings, and the design system. Declare a case on the Cases view writes an uncommitted cases/<id>/case.json into the checkout. The run history shows a run that is executing, with its stage, elapsed time, and spend, and the Live monitor follows it live. New run and Replay a step on run history and Replay from here on a pipeline stage start a run or replay under the stored spend ceiling, which spends provider money. Several prototype screens are still planned. The server is intended for local use, binds to IPv4 loopback, and has no authentication.

Documentation

I want to… Read
Understand the goals and long-term direction Vision
Understand context visibility and efficiency plans Context assessment and roadmap
Know what works and what needs work Current state and priorities
Run an experiment and inspect its evidence Runbook
Configure cases, replay, grading, and comparisons Harness reference
Understand the implementation Architecture
Contribute code or keep the docs current Contributing
Look up a project term Glossary
Read the evaluation methodology Research

The documentation index also explains the status of the UI design and recovered historical material. Public project context lives in these tracked documents; the maintainer's personal task board is outside this repository.

License

Licensed under the Apache License 2.0.

About

Benchmark and improve coding-agent instructions with real Claude Code sessions, stage replay, and graded evidence.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages