A flight recorder for tool-calling LLM agents: 50/50 recorded runs replayed byte-identical, with 5.2% recorder overhead at realistic LLM latency.
- Production agent incidents do not reproduce: prompts changed, tools mutated state, and the failure is gone by the time anyone looks.
- Agent logs are mutable and partial, so post-incident analysis cannot prove what the model actually saw at step N.
- Runaway agents burn steps, tokens, and dollars with no enforced ceiling and no recorded reason for why they stopped.
When a production LLM agent misbehaves, the team usually has nothing to debug with. The prompt template was edited last Tuesday, a tool wrote to a database, retries reordered events, and the incident cannot be re-run. Classic observability answers "how long did it take"; it does not answer "what exactly did the model see, and would it do the same thing again". That gap is what this project closes.
The mechanism is a flight recorder, not a logger. flightrec wraps the two boundaries where non-determinism enters an agent (the LLM adapter and the tool registry) and appends every LLM call, tool call, state transition, and budget decision to an append-only JSONL trace. Each event embeds the previous event's hash, so any edit, deletion, or reorder of the stored trace is detected by flightrec verify, which names the first broken link. Deterministic replay re-runs the agent with all responses served from the trace: if the agent recomputes an input that differs from what was recorded, replay stops at that exact step and prints a unified diff (ADR-002). A budget guard enforces per-run ceilings on steps, tokens, simulated dollars, and wall clock, halting the agent with a typed reason that becomes part of the tamper-evident record.
All numbers below were measured on a 2 vCPU, 4GB shared container with a deterministic scripted LLM adapter (no external APIs; the adapter is a labeled stand-in, the recorder and replay engine are model-agnostic). 50 of 50 recorded runs, including budget-halted and fault-injected runs, replayed to byte-identical final states, and all 50 hash chains verified. Recorder overhead measured 5.18% under simulated 25 ms LLM latency and 761% on a zero-latency stress loop, which is the honest worst case: when the whole agent step takes microseconds, any per-event cost dominates. Trace writes sustained 13,400 hash-chained events/sec at 100,000 events. Test coverage is 94% (pytest-cov).
flowchart LR
subgraph runtime [Agent runtime]
A[Agent step loop] -->|complete| L[LLM adapter]
A -->|call| T[Tool registry]
A -->|check before each step| B[Budget guard]
end
subgraph record [Record path]
R[Recorder] -->|append hash-chained JSONL| S[(Trace store)]
end
L -.every call.-> R
T -.every call.-> R
B -.typed halt reason.-> R
subgraph forensics [Offline forensics]
S --> V[verify: chain integrity]
S --> P[replay engine: serves recorded responses]
S --> D[diff: first divergent step]
P -->|input hash mismatch| D
end
classDef fb stroke:#dc2626,stroke-width:2px
class L,T,B fb
Red-bordered nodes are the failure boundaries: non-determinism and spend can only enter through the LLM adapter, the tool registry, and the clock read by the budget guard, so those are exactly the interception points for both recording and replay. Trace file corruption and chain breaks are caught downstream by verify; replay drift is caught by the replay engine and localized by diff.
| Component | Choice | Why here specifically |
|---|---|---|
| Language | Python 3.10+ | Where agent runtimes live; stdlib hashlib/json cover canonical hashing with zero heavy deps |
| Event schema | pydantic v2 | Every trace line is validated on read and write; a corrupt event fails loudly with the line number |
| CLI | click | Subcommands with real exit codes so verify and replay work in CI gates |
| Config | PyYAML + env vars | Pricing table and redaction keys change per deployment, code should not |
| Hashing | SHA-256 over canonical JSON | Key-order-independent encoding makes hashes stable across processes |
| Tests | pytest + pytest-cov | 55 tests, 94% coverage, coverage gate at 90% in CI |
| Lint | ruff | Single fast tool for lint plus import order in CI |
| Demo agent | Hand-rolled typed step loop | langgraph did not install cleanly in the offline build environment within the time budget, so the demo uses a typed state loop; the recorder wraps adapter boundaries and is framework-agnostic (see flightrec/demo/agent.py) |
| Demo LLM | Scripted deterministic adapter | Labeled stand-in, not a model: emits seeded tool-call plans so record/replay/budget behavior is testable offline with zero API keys |
git clone https://github.com/vedantpanchal/agent-flight-recorder
cd agent-flight-recorder
python3 -m venv venv && . venv/bin/activate
pip install -e ".[dev]"
# 1. Record a demo agent run (scripted LLM, no API keys needed)
flightrec record-demo --scenario refund_flow --seed 7 --out trace.jsonl
# 2. Verify the hash chain
flightrec verify trace.jsonl
# 3. Deterministically replay it: byte-identical final state
flightrec replay trace.jsonl
# 4. Simulate a tool whose behavior changed after recording,
# then watch replay pinpoint the first affected step
flightrec tamper trace.jsonl --step 5 --path output.result \
--value 999999 --out mutated.jsonl --reseal
flightrec replay mutated.jsonl # exits 1, prints unified diff at step 7
# 5. Trip a budget ceiling and see the typed halt in the trace
flightrec record-demo --scenario runaway_loop --seed 3 \
--max-steps 6 --out halted.jsonl
flightrec stats halted.jsonl
# 6. Compare any two traces step by step
flightrec diff trace.jsonl mutated.jsonl
# Run the tests and benchmarks yourself
pytest --cov=flightrec
python benchmark/run_benchmarks.py --quickCaptured outputs of steps 4 and 5 live in benchmark/results/divergence_demo.txt, benchmark/results/tamper_evidence_demo.txt, and benchmark/results/budget_demos/.
Methodology: all runs on a 2 vCPU, 4GB shared container (another build shared the box; batches were repeated and medians reported). The demo LLM is the deterministic scripted adapter, so no network time is included. Overhead compares the identical scripted workload with the recorder writing JSONL versus a no-op recorder, under two conditions: a zero-latency stress loop (worst case: the whole agent step costs microseconds) and simulated 25 ms LLM latency (labeled simulation of a fast production agent). Throughput writes synthetic llm_call events (about 0.9 KB each) through the full hash-chain path. Reproduce with python benchmark/run_benchmarks.py; raw output in benchmark/results/benchmarks.json.
| Trace size | Write events/sec | Write time | File size | Verify events/sec | Chain verified |
|---|---|---|---|---|---|
| 1,000 events | 10,389 | 0.10 s | 0.91 MB | 24,368 | yes |
| 10,000 events | 14,382 | 0.70 s | 9.05 MB | 21,713 | yes |
| 100,000 events | 13,408 | 7.46 s | 90.6 MB | 15,603 | yes |
| Overhead condition | Recorder off (median) | Recorder on (median) | Overhead |
|---|---|---|---|
| Zero-latency stress, 40 runs/batch, 5 repeats | 0.0079 s | 0.0683 s | 761% |
| Simulated 25 ms LLM latency, 10 runs/batch, 3 repeats | 1.0416 s | 1.0956 s | 5.18% |
Honest degradation: verify throughput drops about 36% from 1k to 100k events because verification currently parses and pydantic-validates the entire 90 MB trace in memory before walking the chain, and the zero-latency overhead number is the price of hashing full payloads when the agent itself does almost no work; a real agent spends its wall clock inside model calls, which is what the 25 ms condition models.
Two ADRs in docs/adr/:
- ADR-001: hash-chained JSONL event log over OpenTelemetry spans for agent forensics. Tamper evidence, full payloads for replay, offline diffing; the cost is sensitive trace files and no free tracing UI.
- ADR-002: record/replay at the adapter boundary over deterministic seeding of the whole runtime, with an explicit trigger to revisit if in-process non-determinism between adapter calls becomes material.
- Live token-stream capture. Traces record complete request/response pairs. Trigger to add: a real incident where the divergence happened inside a streamed partial response, proving non-streaming traces insufficient.
- Multi-agent trace stitching. One trace is one agent run. Trigger to add: single-agent forensics adopted in practice and incidents start crossing agent boundaries.
- OTel span export. The JSONL file is the authoritative record; an exporter is future work, not a dependency.
- Signing the chain head with an external key. The chain proves integrity of the file against edits, not authorship; signing matters once traces cross trust boundaries.
- Traces contain full prompts, tool arguments, and outputs by design (replay needs them). Treat the trace store like a database of user data: restrict access, encrypt at rest (filesystem or volume-level encryption; the recorder writes plain JSONL), and set retention.
- Redaction hooks are provided:
Recorder(redactor=make_key_redactor(...))masks configured keys anywhere in payloads before hashing and writing, so secrets never reach disk and verification still passes. Configure keys viaFLIGHTREC_REDACT_KEYS. - Secrets never live in config files: pricing and redaction keys come from env vars (
FLIGHTREC_*) orflightrec.yaml; the library never logs API keys, and the demo needs none. - Structured JSON logs go to stderr and contain run ids and counts, never payloads.
| Failure | Detection | Behavior | Recovery |
|---|---|---|---|
| Trace file corruption mid-write (crash, torn line) | load_trace raises TraceCorruption with the line number |
Verify and replay refuse to run on the corrupt file | load_trace(allow_partial=True) recovers the valid prefix, which still verifies; the tail is lost and says so |
| Hash chain break (edited, deleted, reordered events) | flightrec verify recomputes every hash |
Exit 1 with the first bad step and reason | Treat the trace as evidence of tampering; restore from the write-once store |
| Replay drift from non-deterministic tools | Replay hash-compares every recomputed input | Replay stops at the first divergent step, prints a unified diff, exit 1 | Fix or pin the non-deterministic input; the diff names it |
| Budget guard failure direction | Guard is checked before every step and fails closed: an exception in the guard halts the run rather than letting it spend | Typed budget_halt event with reason, limit, actual; enforcement is pre-step, so a run can overshoot a ceiling by at most one step's consumption |
Raise the ceiling or fix the loop; the trace shows exactly what was consumed |
| Recorder cannot write (disk full, permissions) | record() raises immediately |
Run fails loudly rather than running unrecorded | Fail-open mode is deliberately not offered; an unrecorded agent is the incident |
The first overhead benchmark reported 772% recorder overhead and my first instinct was that the number had to be wrong. Profiling 20,000 events showed it was real and self-inflicted: every event was being serialized four times (input hash, output hash, chain hash, line write) plus two pydantic dumps, 1.38 million function calls for 20k events. I rewrote the hot path to serialize the canonical body once and splice the sealed hash into that same string (safe because parsers do not depend on key order and verification recomputes from the parsed model). Write throughput went from about 7,350 to about 13,000 events/sec (fix commit 6c969dd). The second lesson was the more important one: overhead only moved from 772% to 761% on the stress loop, because the baseline agent step costs microseconds, so the ratio measures the workload, not the recorder. That is why the README reports both the worst-case ratio and the realistic-latency condition (5.18%) instead of one flattering number.
- Streaming verify and replay that walk the JSONL file without loading the full trace into memory (removes the 36% verify degradation at 100k events).
- OTel span exporter emitting redacted spans alongside the authoritative JSONL trace.
- Chain-head signing (age/ssh keys) so traces are authentic, not just tamper-evident.
- Adapter shims for LangGraph and the OpenAI/Anthropic SDK call sites, replacing the scripted demo adapter in real deployments.
flightrec replay --againstto replay one trace's agent against another trace's responses for cross-version debugging.
MIT. Everything in flightrec/ is author-built on the Python standard library plus pydantic, click, and PyYAML; no agent framework code is vendored.