A verifiable local memory architecture for long-running AI agents.
Memory Forest organizes evidence into a numbered filesystem, preserves where claims came from, and returns routes before it returns memory bodies. Structured records stay human-readable in Markdown, while raw chronology can remain bounded JSONL. Local indexes are derived and replaceable.
Important
Memory Forest is not a hosted service, an authorization system, or a promise that an AI agent will remember correctly. Memory is grounding data. The caller must open the routed source, resolve conflicts, check freshness, and apply its own authority and safety rules.
Long-running agents usually fail in one of two ways. They keep too little context, or they accumulate an unstructured pile that is difficult to retrieve and impossible to audit. Memory Forest separates source, lifespan, structure, promotion, and retrieval. Capture, readable evidence, working detail, durable knowledge, and long-horizon anchors remain connected by an inspectable provenance path.
The design is built around these properties.
- Canonical local storage - the operator-selected filesystem owns memory content and provenance, while indexes remain replaceable.
- Route-first retrieval - default queries return relative paths and bounded ranking metadata, not raw private bodies.
- Root-first retrieval trails -
retrieveranks bounded lexical matches, materializes each match through the canonical XLTM, LTM, MTM, and STM ownership chain, and returns a freshly hash-checked trail. - Provenance-preserving promotion - a durable summary keeps a pointer to the evidence that justified it.
- Rebuildable derived state - indexes can be deleted and recreated without changing canonical memory.
- Mechanical boundaries - validators and audits check layer ownership, link direction, path safety, and common public-release mistakes.
The tree is canonical because ownership should be deterministic. Wikilinks preserve the evidence path between adjacent layers. Similarity edges, graph communities, embeddings, and visualizations can still be generated, but they remain disposable derived state instead of silently redefining where a memory belongs.
This connected trail can keep knowledge, dated decisions, explicit responsibilities, projects, and time-bound provenance related without turning those relationships into access authority. Identity and permission policy stay with the integrating application.
flowchart TB
source["Observed source"] --> istm["06 ISTM<br/>raw chronology and provenance"]
istm --> daily["05 Daily<br/>readable source lane"]
daily --> stm["04 STM<br/>detailed dated leaves"]
stm --> mtm["03 MTM<br/>active recurring branches"]
mtm --> ltm["02 LTM<br/>durable trees"]
ltm --> xltm["01 XLTM<br/>forest and long-horizon anchors"]
stm -. archive candidate .-> archive["00 Life Archive<br/>reusable historical record"]
mtm -. archive candidate .-> archive
ltm -. archive candidate .-> archive
xltm --> route["Root-first trail materialization"]
route --> metadata["Relative path metadata"]
metadata --> open["Explicit canonical source open"]
open --> verify["Freshness and conflict check"]
Capture moves upward from recent evidence toward durable structure. The v0.2 retrieve command globally ranks bounded lexical evidence, then materializes each candidate in root-first canonical ownership order. It reopens the selected files, verifies their indexed hashes, and returns metadata-only trails. This is not a staged top-down semantic traversal. An XLTM-only lexical match remains a depth-one partial trail instead of fanning out across the whole forest. The existing route and search commands retain their v0.1 flat-query behavior and JSON boundary. 00 life_archive is a side archive for reusable history, not a higher truth rank.
See the end-to-end retrieval guide, Architecture, and Layer contracts for the full model.
| Layer | Object | Responsibility |
|---|---|---|
00 life_archive |
archive record | reusable project or life history with retained provenance |
01 xltm |
forest | long-horizon anchors and cross-tree invariants |
02 ltm |
tree | durable knowledge and stable thematic contracts |
03 mtm |
branch | active recurring work, procedures, and medium-horizon state |
04 stm |
leaf | detailed dated evidence and short-term actionable context |
05 daily |
daily source | readable, source-bound digest of recent events |
06 istm |
raw source | append-oriented chronology and provenance records |
Layer placement is a responsibility decision, not a confidence score. A claim in XLTM can still be stale or wrong.
Memory Forest is currently an alpha reference implementation for macOS and Linux. It requires Python 3.11 or newer, a POSIX filesystem with Unix permission modes, and no required runtime dependencies outside the standard library.
git clone https://github.com/hyungchulc/memory-forest.git
cd memory-forest
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .Create a private synthetic demo with strict runtime permissions, then inspect and query it.
init accepts only a new path whose direct parent already exists as a real, non-symlink directory. It never creates missing ancestor directories and refuses every existing target.
demo_parent="$(mktemp -d)"
demo_root="$(cd "$demo_parent" && pwd -P)/forest"
memory-forest init "$demo_root" --example
memory-forest doctor "$demo_root"
memory-forest validate "$demo_root"
memory-forest audit "$demo_root"
memory-forest index "$demo_root"
memory-forest route "$demo_root" "instrument calibration"
memory-forest search "$demo_root" "reference lamp"
memory-forest retrieve "$demo_root" "telemetry replay"Search is metadata-only by default and queries the last built index snapshot. Reading matching bodies is an explicit operation. When a selected canonical file has changed since indexing, body retrieval fails with index_stale instead of returning cached text. Rebuild the index after forest changes.
memory-forest search "$demo_root" "reference lamp" --include-bodyThe local core does not call a translation, embedding, or model API. A caller may supply a strict QueryPlan containing query strings only. This allows an OAuth-authenticated gateway to expand a Korean, English, mixed-language, or other Unicode query while keeping tokens, filesystem roots, memory bodies, and authorization outside the plan.
{
"schema_version": 1,
"probes": [
{"query": "mission recovery"},
{"query": "telemetry replay"}
]
}printf '%s' '{"schema_version":1,"probes":[{"query":"mission recovery"}]}' \
| memory-forest retrieve "$demo_root" "비상 복원" --query-plan -Every probe object must contain exactly one query field. Paths, bodies, tokens, credentials, provider settings, and all other fields are rejected. The original query is always retained. A trail with direct original-query evidence always ranks ahead of a plan-only trail; accepted probes add recall and ranking evidence only within that boundary. The expansion quality is owned by the caller; SQLite Unicode tokenization and supplied probes do not guarantee semantic retrieval for every language.
See OAuth and API integration, the QueryPlan schema, and the CLI reference.
The tracked synthetic forest is a human-readable reference. Git does not preserve the private 0700 directory and 0600 file modes required of a runnable forest, so use init --example rather than validating the checkout fixture in place. pwd -P resolves platform temp-directory aliases before the CLI applies its no-symlink boundary. Remove the temporary demo when you are finished.
See the CLI reference before integrating the output into another process.
v0.3 provides three network-free standard-library write commands. The
integrated Structured path is the reference path; promote remains a bounded
leaf-promotion compatibility command.
chmod 600 daily-plan.json structured-sweep-plan.json
memory-forest apply-daily "$demo_root" daily-plan.json
memory-forest apply-structured "$demo_root" structured-sweep-plan.jsonapply-daily writes only the canonical dated Daily source.
structured-context opens bounded, hash-bound current XLTM/LTM/MTM/STM bodies
for one integrated review. apply-structured accepts semantic targets across
all four structured layers, applies exact creates or full-body replacements in
one transaction, and refreshes validation, audit, and the derived index once.
The semantic hierarchy is exactly XLTM Forest, LTM Tree, MTM Branch, and STM
Leaf. Structured targets use tree as the LTM routing key; it is not a
separate object level.
The write commands acquire the same sibling maintenance lock, reject symlink
and case-fold ambiguity, roll back handled validation/audit/index failures, and
publish a private receipt only after the new index succeeds.
Each plan is bound to the stable private forest_id created by init, so
replacing a forest at the same pathname fails closed. Empty reviewed arrays
close as receipt-backed no-ops.
Legacy schema-v1 forests without forest_id remain valid for read-only
validation and retrieval. The v0.3 writers reject them until a supported
migration assigns an identity; they never edit legacy configuration implicitly.
See Integrated Structured sweep, Daily Plan v1, Structured Sweep Plan v1, Write Receipt v1, and the CLI reference.
The default retrieval boundary is deliberately narrow.
- A query is evaluated against an exact local forest root and its private derived index.
- The tool returns bounded route metadata such as layer, relative path, title, and score.
retrievereopens only the selected trail files to verify current hashes, then discards their bodies before emitting JSON.- The caller separately chooses whether to open a canonical body.
- Any opened body is handled under the caller's own model, account, retention, and privacy controls.
The route result is not source truth and does not grant permission. A local SQLite index should be protected like the source forest because it may contain derived text or tokens.
Read Privacy and trust, PRIVACY.md, and SECURITY.md before connecting a real forest to an agent.
Memory Forest Retrieve is a separate caller-owned integration profile for assistants that must consult memory before responding to every non-empty user-authored text turn.
Its example gate runs both route and retrieve, returns a metadata-only current-turn receipt, and treats zero matches as a successful lookup with no evidence. The prompt and companion skill are advisory. Hard enforcement requires the host to register the gate before response generation and block normal completion without a successful receipt for that exact turn.
The gate never auto-indexes, repairs, scans another root, returns bodies, or uses the network. Route and retrieve metadata remain private and untrusted.
Memory Forest owns the canonical layer and retrieval contracts. Source collection and retrieval evaluation are separated into focused companion projects.
- Codex Context for ISTM incrementally collects local Codex conversation sessions into ISTM and produces bounded Daily digests with launchd examples on macOS.
- Mac Context for ISTM collects local Apple Mail, Notification Center, Reminders, and Calendar context into a separate private ISTM store on macOS.
- Memory Retrieval Lab measures retrieval quality with synthetic fixtures, reproducible metrics, multilingual cases, and ranker adapters.
The reference environment used GPT-5.6 Sol with reasoning effort xhigh on 2026-07-24. The collectors and deterministic baselines are model-independent. Read Daily and ISTM companion projects for the exact repository boundaries and bounded Daily contract.
Promotion is an adjacent-layer, evidence-preserving operation.
06 ISTM -> 05 Daily -> 04 STM -> 03 MTM -> 02 LTM -> 01 XLTM
\-------- 00 Life Archive candidate --------/
A promoted record should carry the source pointer, source and capture times, scope, observed fact, derived conclusion, uncertainty, destination, and reason. Promotion does not make a claim true. Mutable facts still require current verification.
The Life Archive arrows above represent selection from structured history, not unrestricted canonical wikilinks. In the current schema, Life Archive links canonically only to adjacent XLTM; nonadjacent evidence remains a plain provenance path.
The reference workflow makes one integrated Structured decision over current
XLTM/LTM/MTM/STM plus committed Daily. Parent-before-child is the internal
materialization order within that sweep, not a separate parent-first workflow.
audit requires every LTM, MTM, and STM record to link to its immediate
canonical parent. Same-layer lateral wikilinks are rejected so ownership
remains mechanically inspectable.
Memory Forest can be validated and reindexed with POSIX cron, a macOS LaunchAgent, or a Codex Scheduled Task. Reviewed plans can also be applied locally with the same sibling maintenance lock.
The core does not collect source systems or decide semantic promotions. The automation starter therefore separates two lanes:
- implemented deterministic maintenance, which locks one exact private root, validates and audits it, then atomically rebuilds the derived index
- caller-owned integrated layer and structure judgment, followed by the
implemented
apply-dailyorapply-structuredtransaction writer with provenance binding, rollback, validation, audit, index, and receipt proof
The repository includes a shared maintenance wrapper, a crontab example, a per-user macOS launchd template, and a bounded Codex Scheduled Task prompt. Read the Automation guide before enabling unattended runs. Cron and launchd can remain local. A Codex task uses the selected account and model processing boundary, so do not send strictly local-only evidence through that lane.
Codex Debug Bridge includes a small memory-forest-starter as an onboarding path for a private route-only helper. This repository is the standalone portable reference implementation. It adds the CLI, contracts, audits, synthetic example, documentation, and companion skill.
Neither project contains a real private forest. Neither project includes private prompts, unattended source collection or semantic-plan generation, operational logs, personal adapters, or user identifiers. The projects can be used independently.
See Codex Debug Bridge integration.
Memory Forest is alpha software. The v0.3 scope includes strict receipt-backed Daily application, bounded current-Forest context, one integrated XLTM/LTM/MTM/STM sweep, deterministic root-first retrieval, portable contracts, synthetic fixtures, and a route-first body boundary. It is suitable for evaluation and adaptation, not unattended high-stakes decision-making.
Candidate directions include measured multilingual routing evaluation, optional ranking adapters, generated graph views, migration tooling, and stronger provenance integrity checks. These are not release promises. See Roadmap.
| Path | Purpose |
|---|---|
src/memory_forest/ |
local CLI and reference implementation |
docs/ |
architecture, contracts, privacy boundary, and integration guidance |
docs/retrieval-guide.md |
query intake, deterministic retrieval, explicit body boundary, and integration limits |
docs/daily-and-istm-companions.md |
bounded Daily contract and companion repository boundaries |
examples/synthetic-forest/ |
fictional 00-06 forest with no private identifiers |
examples/automation/ |
cron, launchd, and Codex Scheduled Task maintenance examples |
examples/memory-forest-retrieve/ |
metadata-only mandatory consultation gate and prompt |
skills/memory-forest/ |
companion Codex skill |
skills/memory-forest-retrieve/ |
opt-in always-on retrieval integration skill |
tests/ |
behavior and regression tests |
scripts/audit_public_release.py |
public-tree secret and identifier audit |
Read CONTRIBUTING.md. Keep fixtures synthetic, keep real forests outside Git, and run the full local check before opening a pull request.
make checkSecurity issues should follow SECURITY.md. Community participation follows CODE_OF_CONDUCT.md.
Memory Forest is released under the MIT License.
