PARM is a causal benchmark suite for output-conditioned personal memory in AI agents.
The benchmark targets a difficult and consequential failure mode: an agent begins with an ordinary request, then discovers a decisive cue inside a tool result, document, or other late observation. Prompt-time retrieval cannot see that cue. Naively searching the entire observation can find the right memory, but often floods the agent with unrelated personal context.
PARM asks a stricter question:
Can an agent recognize when a late observation makes a specific memory actionable, retrieve it before the governed action, and stay silent when the cue is removed?
Every scenario is a causal triplet:
- Positive: a source-grounded late cue makes a memory relevant.
- Cue-ablated control: the decisive cue is removed while the rest of the task remains comparable.
- Memory-included ceiling: the agent receives the relevant source directly, verifying that the intended action is achievable.
PARM is built as research infrastructure rather than a one-off prompt demo:
- 157 frozen, independently constructed base scenarios / 471 executable cases across 140 persona-isolated histories in the current PARMBench v1 calibration batch.
- Raw-history, retrieval-agnostic evaluation: PARMBench provides source material and a causal contract, not a preferred memory store or retrieval implementation.
- Evidence and fairness gates: scenario construction checks source support, prompt opacity, persona isolation, cue selectivity, and ceiling actionability before a case can support a claim.
- Executable workflow evaluation: PARMBench Workflows evaluates multi-step tool trajectories from final environment state, timely memory admission, and intervention restraint—without using an LLM judge as the scorer.
- Replayable artifacts: data, manifests, validation profiles, retrieval artifacts, and result sidecars are versioned so claims can be inspected and reproduced.
The benchmark is designed to make progress legible: a retrieval policy must improve the positive case and preserve the cue-ablated control. A system that retrieves aggressively and changes behavior everywhere does not pass.
The current v1 batch contains 157 base scenarios, 471 cases, and 140 persona-isolated histories. Its final frozen replay passes 157/157 scenario contracts. The batch is intentionally retrieval-agnostic: it is a durable, auditable substrate for comparing memory systems fairly rather than an evaluation designed around PARM's own implementation.
See the PARMBench v1 dataset record and construction contract.
On the repaired 54-case Amara development suite, PARM V5 achieved 30/36 positive/control decisions. The strongest comparison conditions achieved 18/36. PARM admitted 15 gold sources with 6 spurious admissions; all-entity output RAG admitted 16 gold sources but 1,361 spurious sources.
| Condition | Correct positive + control decisions |
|---|---|
| PARM V5 | 30/36 |
| No memory | 18/36 |
| All-entity output RAG | 18/36 |
| Naive output RAG | 18/36 |
| Prompted memory-tool agent | 17/36 |
| Enhanced input RAG | 15/36 |
This is an encouraging mechanism result, not a held-out product claim: V5 was improved using the expansion cases. The immutable first pass and the full limitations are preserved in the V5 report.
The workflow suite currently validates 18 deterministic cases across tool-using agent environments. It measures the state an agent leaves behind, the sources it admits, and whether memory arrives before the action it is meant to govern. This closes an important gap between one-shot choice evaluation and realistic agent execution.
The workflow work is deliberately candid: it records successful interventions, false interventions, ceiling failures, and retired candidates. That discipline is a feature—new workflow families are certified before their policy results are used as evidence. Start with the workflow suite, scenario-set result, and applicable-memory handoff experiment.
python -m venv .venv
. .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install -e .
PYTHONPATH=src python -m parm_bench.cli validate data/benchmark_parmbench_v1
PYTHONPATH=src python -m parm_bench.cli workflow validate data/workflows_v1
PYTHONPATH=src python -m unittest discover -s testsLive model-backed replay requires OPENAI_API_KEY in a local, ignored .env; canonical validation and deterministic scoring do not.
src/parm_bench/ benchmark package and CLI
src/parm_bench/workflows/ environments, runners, policies, and verifiers
tests/ unit and CLI coverage
data/benchmark_parmbench_v1/ frozen raw-history causal calibration batch
data/workflows_v1/ executable multi-step workflow scenarios
data/retrieval-indexes/ frozen retrieval artifacts
data/benchmark-results/ replayable predictions, configs, and metrics
docs/ claim contracts, architecture, evidence, and roadmap
| Document | What it covers |
|---|---|
| Research claim and scope | The late-cue retrieval problem, hypotheses, and explicit boundaries |
| Architecture | Retrieval, admission, handoff, and workflow components |
| Evaluation contract | Causal triplets, deterministic metrics, and failure taxonomy |
| Scenario construction | Evidence, isolation, fairness, and ceiling gates for new cases |
| PARMBench v1 calibration batch | 157-scenario raw-history benchmark record |
| Workflow suite | Multi-step agent environments and final-state verification |
| V5 development result | Selective-retrieval comparison, artifacts, and limitations |
| Workflow scenario set | Cross-scenario workflow results and construction lessons |
| Roadmap | Next evaluation and external-validity work |
PARM is ambitious about the problem and conservative about the evidence. It does not claim to have solved general long-term memory, broad agent reliability, or real-world personal-agent usefulness. It provides a rigorous way to test one concrete capability that current evaluations frequently blur: whether memory should intervene only after the environment gives it a reason to.