Local MCP server that makes development-loop reports token-efficient for LLM coding agents — profiler output first, and since v1.2 the rest of the loop's feedback artifacts too (build traces, build diagnostics, benchmark results, CI run logs, repo change state). It sits between the tool and the agent: reads the report from disk and returns a small, structured, numeric digest the agent can act on, while keeping a pointer back to the raw report for lazy expansion.
It is a translator/router, not a judge — interpretation ("is this kernel memory-bound?") is the model's job. perfdigest only provides efficient, deterministic access to clean numeric metrics across vendors and languages.
perfdigest deliberately splits what an agent does with a profiler into two tiers:
- Digest (read a report) — universal, available on every machine. A report's
origin is irrelevant to digesting it. An NVIDIA
.ncu-rep(or its CSV export) captured on a CI/remote GPU runner can be pulled to a Mac and digested there — that is how you do CUDA work on a Mac via CI. Tier-1 is never gated by local hardware, only by whether a reader's dependency imports. - Capture (produce a report) — platform-verified. You cannot run
ncuon a Mac, Metal on Linux, or hardware PMU counters under WSL2.platform_capabilitiesandsuggest_profile_commandgate this so the agent never spends context on a capture that can't run here, and redirect it to "capture elsewhere, digest here."
Every change a session makes propagates through the same four stages, each with
its own verbose artifact: what changed (RepoState) → does it build, at what
cost (BuildDigest) → does it pass (CIDigest) → how fast is it (PerfDigest).
perfdigest digests all four through the same seven tools — a backend is a
registry row, not a new API surface. One principle binds the whole matrix: we
bind to artifacts, never to live APIs — every reader parses a file on disk
whose format is frozen by its ecosystem's backward-compat pressure. (CI
artifacts need no dedicated backend at all: a downloaded .ncu-rep from a CI
runner is digested by nsight — capture anywhere, digest anywhere, by
composition.)
| Channel | Backend | format |
Domain | Capture tool | Digest anywhere |
|---|---|---|---|---|---|
| Perf | nsight |
ncu-rep |
gpu_kernel | NVIDIA ncu (Linux/Windows) |
needs ncu-report wheel |
| Perf | nsight_csv |
ncu-csv |
gpu_kernel | (export of ncu) |
✅ pure Python |
| Perf | rocm |
rocprof-csv |
gpu_kernel | AMD rocprof (Linux/Windows) |
✅ pure Python |
| Perf | linux_perf |
perf-stat-json, perf-report |
cpu_function | Linux perf |
✅ pure Python |
| Perf | metal |
metal-trace |
gpu_pass | Apple xctrace (macOS) |
✅ pure Python |
| Perf | chrome_trace |
torch-trace, chrome-trace |
framework_op + gpu_kernel | torch.profiler → export_chrome_trace() |
✅ pure Python |
| Build | ptxas |
ptxas-verbose |
kernel_codegen | nvcc -Xptxas -v (no GPU needed) |
✅ pure Python |
| Build | clang_time_trace |
clang-time-trace, ftime-trace |
build_phase | clang -ftime-trace |
✅ pure Python |
| Build | cmake_profile |
cmake-profile, cmake-trace |
build_phase | cmake --profiling-format=google-trace (≥3.18) |
✅ pure Python |
| Build | ninja_log |
ninja-log, ninja_log |
build_step | .ninja_log (byproduct of any ninja build) |
✅ pure Python |
| Build | cargo_diag |
cargo-diag, cargo-json |
build_diag | cargo build --message-format=json |
✅ pure Python |
| Build | criterion |
criterion, criterion-json |
benchmark | cargo bench (target/criterion dir) |
✅ pure Python |
| CI | gha_log |
gha-log, gh-run-log |
ci_step | saved gh run view <id> --log |
✅ pure Python |
| RepoState | git_numstat |
git-numstat, numstat |
repo_change | saved git diff --numstat -M |
✅ pure Python |
GPU backends share one vocabulary (compute_pct_peak, dram_pct_peak,
l2_hit_rate, achieved_occupancy, …); the CPU backend introduces a CPU
vocabulary (ipc, cache_miss_rate, llc_miss_rate, branch_mispredict_rate,
self_pct, …). The ptxas backend adds the codegen layer: runtime counters
say what is slow, registers_per_thread / spill_stores_bytes /
static_smem_bytes say why the code became what it is — captured at compile
time, so it works in any CI job with no GPU attached.
RepoState is deliberately session-retrospective, not GitHub-shaped: a saved
numstat snapshot is the token-cheap answer to "what has this session changed
so far", and it doubles as the provenance anchor that pins perf/build/CI reports
to a code state. It emits facts (files, line counts, honest None for binary
files) — never verdicts like "release ready"; concluding is the agent's job.
Future: Go, Java, and other perf-critical runtimes.
- Suppress profiler stdout — write to a file. e.g.
ncu --set full -o report.ncu-rep ./app. If the profiler prints its summary to stdout, that raw table enters the agent's context before perfdigest runs — defeating the entire purpose. Nonemeans "not measured in this export", NEVER zero. A metric the export does not contain is returned asnot_available_in_this_export, not a fake0.0. A genuine0.0(e.g. zero branch divergence) is preserved. Silently returning zero = lying to the model = the worst bug this tool can have.
Tier 1 — digest (any backend, any host):
summarize_report(report_ref, format, top_n=5)→ the N hottest units + core metrics in one call (ranked byduration_us, falling back toself_pct, then file order; reports duration coverage of the returned units)list_kernels(report_ref, format)→[{name, index, duration_us, domain}]get_metrics(report_ref, format, kernel, metrics=None)→ compact digest (metrics=None→ the backend's default core set)compare_metrics(report_a, report_b, format, kernel, kernel_b=None)→ the measure→edit→measure tool:{a, b, delta, delta_pct}per metric withdelta = b − a; a metric missing on either side yields an honestnot_available_in_this_exportdelta, never a fake0.0expand(report_ref, format, kernel, section)→ raw vendor metrics (the safety valve)
Reports are parsed once per file version (an mtime/size-keyed cache), so repeated digests of the same report cost no re-parse.
Tier 2 — capture advisory (platform-verified):
platform_capabilities()→ machine identity +can_digest(universal) vscan_capture_here(gated)suggest_profile_command(backend, target)→ the correct, platform-aware invocation, or a refusal that redirects to the tier-1 path
format is mandatory — the agent passes what it produced; a path says where
a file is, not what format it is.
uvx perfdigest-mcp # run from PyPI (downloadable)
uv tool install "perfdigest-mcp[nvidia]" # + NVIDIA native binary reader (Linux/Windows)PyPI/install name is
perfdigest-mcp; the command and import package areperfdigest(e.g.uvx perfdigest-mcp,import perfdigest).
Claude Code and OpenAI Codex setup (both stdio MCP): see docs/clients.md.
v1.2.0 — the Development Observatory release
(Python 3.10–3.14 incl. free-threaded 3.14t, CI-verified on Linux/macOS/Windows;
one known upstream gap: Windows 3.14t cannot install yet because mcp's
pywin32 dependency ships no cp314t wheel — that CI cell is a documented
allowed-failure until pywin32 catches up):
seven new backends extend the digest matrix from profilers to the whole
edit→build→verify→measure loop. BuildDigest: clang_time_trace (compiler-phase
traces, with [total] aggregate tagging), cmake_profile (configure-step
traces — cmake emits bare B/E event pairs, stack-folded honestly),
ninja_log (per-edge build timings, v5–v7 logs), cargo_diag (rustc
diagnostics — a backend whose durations are honestly absent), criterion
(Rust benchmark estimates; first directory-shaped report ref). CIDigest:
gha_log (saved gh run view --log files, error/warning annotations,
previous-green comparison proven on real two-run fixtures). RepoState:
git_numstat (session-retrospective change digest; binary files surface as
honest absence, never 0.0). Zero new tools — every backend rides the same
seven; 206 tests. Every fixture in the release was produced by the real tool it
mimics — several briefs were corrected by what the artifacts actually contained.
v1.1.1 — contract-hardening hotfix (issues #2/#4/#5): registry format
collisions are detected on the normalized key (a case-variant format can no
longer silently hijack a backend); metrics=[] now means an empty request
(only None selects the default core set); exact unit names win over numeric
index interpretation, with #3 as the explicit index form; unreadable backends
surface a diagnostic in platform_capabilities (can_digest_issues) instead
of a silent False; and publishing is gated — a tag only reaches PyPI after
the full suite, a clean-venv wheel install, and a tool-surface check pass.
v1.1.0 — the agent-loop release: compare_metrics (before/after deltas),
summarize_report (one-call top-N), a parse-once report cache, the ptxas
codegen backend (compile-time registers/spills/smem; capture needs nvcc, not a
GPU), and the chrome_trace backend (torch/Kineto and any Chrome-trace emitter:
framework ops + the CUDA kernels they launched, one report). Also normalizes
perf scope-modifier event names (cycles:u on perf_event_paranoid>=2
hosts).
v1.0.0 — multi-backend registry; NVIDIA (native + CSV), AMD HIP, Linux perf (C++/Rust), and Apple Metal adapters; platform capability gating; cross-client config. Validated on the Linux/macOS/Windows CI matrix (pure-Python readers run hardware-free against committed fixtures). Real binary-capture tests are fixture-gated and skip without the device.
Every number here comes from a real capture on real hardware, published with the
workload that produced it in
PerfDigest-MCP-Bench
(local eval/README.md is the pointer). Ratios are digest
payload vs the raw report an agent would otherwise read.
| Study | Subject | What it measured |
|---|---|---|
| Demanding workloads pt. 1 | closed-source cuBLAS GEMMs; torch.compile inference | four vendor kernel families diagnosed from counters alone — no source to read — at 29–46×; Inductor fusion visible structurally at 74–115× |
| pt. 2 | a full torch.compile training step (forward + backward + AdamW) |
Inductor-fused backward kernels diagnosed memory-bound at 96–450×; and a compare_metrics pairing caught reporting the wrong sign of a fusion win when one kernel is compared against what is really its multi-kernel replacement |
| pt. 3 | one C++ project through every build lens (cmake_profile + ninja_log + clang_time_trace) |
a template-heavy TU dominating the build 5× over the next edge, its cost split by -ftime-trace into a single recursive-constexpr evaluation (37% of compile time) and STL instantiation pressure (10.7× nesting overlap, which is why aggregate spans are tagged); 261–5,440× on the trace lenses |
| pt. 4 | the CUDA optimization loop (naive → coalesced → vectorized) | v0→v1 a real 2.64× win explained by dram_pct_peak 36%→90%; v1→v2 an honest non-win — a genuine −20.9% instruction cut bought nothing once already at ~90% of DRAM peak; ptxas foreshadows the ranking only below the roofline ceiling; 11–86× |
| Interpreter matrix | six CPythons, 3.10 → free-threaded 3.14t | digests byte-identical across all six for 10 of 11 pure-Python backends (the one divergence traced to CPython 3.12's sum() change); readers parse in < 2 ms; the parse-once cache holds correctness under 8-thread free-threaded contention at 2.36× GIL throughput |
The low ratios are published next to the high ones on purpose. A .ninja_log digests
at 0.80× and its summarize_report at 0.49× — a build log is already terse,
so there is nothing to compress; what the digest buys there is a uniform contract,
not fewer tokens. A measurement tool that only reported its wins would not be one.
perfdigest was authored independently; these adjacent MCP servers occupy a nearby space and likely work well for their narrower scope. The differences are the reason perfdigest exists:
- nsys-mcp (NVIDIA Nsight Systems) — profiles binaries and aggregates trace timeline stats. perfdigest targets Nsight Compute per-kernel counters, is read-only (does not run the profiler), and spans multiple vendors.
- pprof-analyzer-mcp / Profiler-MCP — Go (and Go/Python/Java) CPU/memory profiles, often rendering flamegraphs. perfdigest is a numeric digester focused on token/context efficiency rather than visualization.
What is distinct here: the multi-backend matrix (NVIDIA + AMD HIP + CPU perf +
Metal under one neutral contract), the token-efficiency thesis with the
None≠0.0 honesty rule, and the read/capture split with platform capability
gating (digest anywhere, capture only where supported).