Skip to content

hyungchulc/memory-forest

Repository files navigation

Memory Forest

A verifiable local memory architecture for long-running AI agents.

한국어 README

CI Python 3.11+ License: MIT Local-first

Memory Forest organizes evidence into a numbered filesystem, preserves where claims came from, and returns routes before it returns memory bodies. Structured records stay human-readable in Markdown, while raw chronology can remain bounded JSONL. Local indexes are derived and replaceable.

Important

Memory Forest is not a hosted service, an authorization system, or a promise that an AI agent will remember correctly. Memory is grounding data. The caller must open the routed source, resolve conflicts, check freshness, and apply its own authority and safety rules.

Why this exists

Long-running agents usually fail in one of two ways. They keep too little context, or they accumulate an unstructured pile that is difficult to retrieve and impossible to audit. Memory Forest separates source, lifespan, structure, promotion, and retrieval. Capture, readable evidence, working detail, durable knowledge, and long-horizon anchors remain connected by an inspectable provenance path.

The design is built around these properties.

  • Canonical local storage - the operator-selected filesystem owns memory content and provenance, while indexes remain replaceable.
  • Route-first retrieval - default queries return relative paths and bounded ranking metadata, not raw private bodies.
  • Root-first retrieval trails - retrieve ranks bounded lexical matches, materializes each match through the canonical XLTM, LTM, MTM, and STM ownership chain, and returns a freshly hash-checked trail.
  • Provenance-preserving promotion - a durable summary keeps a pointer to the evidence that justified it.
  • Rebuildable derived state - indexes can be deleted and recreated without changing canonical memory.
  • Mechanical boundaries - validators and audits check layer ownership, link direction, path safety, and common public-release mistakes.

The tree is canonical because ownership should be deterministic. Wikilinks preserve the evidence path between adjacent layers. Similarity edges, graph communities, embeddings, and visualizations can still be generated, but they remain disposable derived state instead of silently redefining where a memory belongs.

This connected trail can keep knowledge, dated decisions, explicit responsibilities, projects, and time-bound provenance related without turning those relationships into access authority. Identity and permission policy stay with the integrating application.

Architecture

flowchart TB
    source["Observed source"] --> istm["06 ISTM<br/>raw chronology and provenance"]
    istm --> daily["05 Daily<br/>readable source lane"]
    daily --> stm["04 STM<br/>detailed dated leaves"]
    stm --> mtm["03 MTM<br/>active recurring branches"]
    mtm --> ltm["02 LTM<br/>durable trees"]
    ltm --> xltm["01 XLTM<br/>forest and long-horizon anchors"]
    stm -. archive candidate .-> archive["00 Life Archive<br/>reusable historical record"]
    mtm -. archive candidate .-> archive
    ltm -. archive candidate .-> archive

    xltm --> route["Root-first trail materialization"]
    route --> metadata["Relative path metadata"]
    metadata --> open["Explicit canonical source open"]
    open --> verify["Freshness and conflict check"]
Loading

Capture moves upward from recent evidence toward durable structure. The v0.2 retrieve command globally ranks bounded lexical evidence, then materializes each candidate in root-first canonical ownership order. It reopens the selected files, verifies their indexed hashes, and returns metadata-only trails. This is not a staged top-down semantic traversal. An XLTM-only lexical match remains a depth-one partial trail instead of fanning out across the whole forest. The existing route and search commands retain their v0.1 flat-query behavior and JSON boundary. 00 life_archive is a side archive for reusable history, not a higher truth rank.

Memory Forest graph

Fictional root-first retrieval trail

See the end-to-end retrieval guide, Architecture, and Layer contracts for the full model.

Layer map

Layer Object Responsibility
00 life_archive archive record reusable project or life history with retained provenance
01 xltm forest long-horizon anchors and cross-tree invariants
02 ltm tree durable knowledge and stable thematic contracts
03 mtm branch active recurring work, procedures, and medium-horizon state
04 stm leaf detailed dated evidence and short-term actionable context
05 daily daily source readable, source-bound digest of recent events
06 istm raw source append-oriented chronology and provenance records

Layer placement is a responsibility decision, not a confidence score. A claim in XLTM can still be stale or wrong.

Quick start

Memory Forest is currently an alpha reference implementation for macOS and Linux. It requires Python 3.11 or newer, a POSIX filesystem with Unix permission modes, and no required runtime dependencies outside the standard library.

git clone https://github.com/hyungchulc/memory-forest.git
cd memory-forest
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .

Create a private synthetic demo with strict runtime permissions, then inspect and query it.

init accepts only a new path whose direct parent already exists as a real, non-symlink directory. It never creates missing ancestor directories and refuses every existing target.

demo_parent="$(mktemp -d)"
demo_root="$(cd "$demo_parent" && pwd -P)/forest"
memory-forest init "$demo_root" --example
memory-forest doctor "$demo_root"
memory-forest validate "$demo_root"
memory-forest audit "$demo_root"
memory-forest index "$demo_root"
memory-forest route "$demo_root" "instrument calibration"
memory-forest search "$demo_root" "reference lamp"
memory-forest retrieve "$demo_root" "telemetry replay"

Search is metadata-only by default and queries the last built index snapshot. Reading matching bodies is an explicit operation. When a selected canonical file has changed since indexing, body retrieval fails with index_stale instead of returning cached text. Rebuild the index after forest changes.

memory-forest search "$demo_root" "reference lamp" --include-body

Query expansion without putting a model in the core

The local core does not call a translation, embedding, or model API. A caller may supply a strict QueryPlan containing query strings only. This allows an OAuth-authenticated gateway to expand a Korean, English, mixed-language, or other Unicode query while keeping tokens, filesystem roots, memory bodies, and authorization outside the plan.

{
  "schema_version": 1,
  "probes": [
    {"query": "mission recovery"},
    {"query": "telemetry replay"}
  ]
}
printf '%s' '{"schema_version":1,"probes":[{"query":"mission recovery"}]}' \
  | memory-forest retrieve "$demo_root" "비상 복원" --query-plan -

Every probe object must contain exactly one query field. Paths, bodies, tokens, credentials, provider settings, and all other fields are rejected. The original query is always retained. A trail with direct original-query evidence always ranks ahead of a plan-only trail; accepted probes add recall and ranking evidence only within that boundary. The expansion quality is owned by the caller; SQLite Unicode tokenization and supplied probes do not guarantee semantic retrieval for every language.

See OAuth and API integration, the QueryPlan schema, and the CLI reference.

The tracked synthetic forest is a human-readable reference. Git does not preserve the private 0700 directory and 0600 file modes required of a runnable forest, so use init --example rather than validating the checkout fixture in place. pwd -P resolves platform temp-directory aliases before the CLI applies its no-symlink boundary. Remove the temporary demo when you are finished.

See the CLI reference before integrating the output into another process.

Provenance-bound local writes

v0.3 provides three network-free standard-library write commands. The integrated Structured path is the reference path; promote remains a bounded leaf-promotion compatibility command.

chmod 600 daily-plan.json structured-sweep-plan.json
memory-forest apply-daily "$demo_root" daily-plan.json
memory-forest apply-structured "$demo_root" structured-sweep-plan.json

apply-daily writes only the canonical dated Daily source. structured-context opens bounded, hash-bound current XLTM/LTM/MTM/STM bodies for one integrated review. apply-structured accepts semantic targets across all four structured layers, applies exact creates or full-body replacements in one transaction, and refreshes validation, audit, and the derived index once. The semantic hierarchy is exactly XLTM Forest, LTM Tree, MTM Branch, and STM Leaf. Structured targets use tree as the LTM routing key; it is not a separate object level. The write commands acquire the same sibling maintenance lock, reject symlink and case-fold ambiguity, roll back handled validation/audit/index failures, and publish a private receipt only after the new index succeeds. Each plan is bound to the stable private forest_id created by init, so replacing a forest at the same pathname fails closed. Empty reviewed arrays close as receipt-backed no-ops.

Legacy schema-v1 forests without forest_id remain valid for read-only validation and retrieval. The v0.3 writers reject them until a supported migration assigns an identity; they never edit legacy configuration implicitly.

See Integrated Structured sweep, Daily Plan v1, Structured Sweep Plan v1, Write Receipt v1, and the CLI reference.

Route-only privacy boundary

The default retrieval boundary is deliberately narrow.

  1. A query is evaluated against an exact local forest root and its private derived index.
  2. The tool returns bounded route metadata such as layer, relative path, title, and score.
  3. retrieve reopens only the selected trail files to verify current hashes, then discards their bodies before emitting JSON.
  4. The caller separately chooses whether to open a canonical body.
  5. Any opened body is handled under the caller's own model, account, retention, and privacy controls.

The route result is not source truth and does not grant permission. A local SQLite index should be protected like the source forest because it may contain derived text or tokens.

Read Privacy and trust, PRIVACY.md, and SECURITY.md before connecting a real forest to an agent.

Required per-turn retrieval

Memory Forest Retrieve is a separate caller-owned integration profile for assistants that must consult memory before responding to every non-empty user-authored text turn.

Its example gate runs both route and retrieve, returns a metadata-only current-turn receipt, and treats zero matches as a successful lookup with no evidence. The prompt and companion skill are advisory. Hard enforcement requires the host to register the gate before response generation and block normal completion without a successful receipt for that exact turn.

The gate never auto-indexes, repairs, scans another root, returns bodies, or uses the network. Route and retrieve metadata remain private and untrusted.

Companion projects

Memory Forest owns the canonical layer and retrieval contracts. Source collection and retrieval evaluation are separated into focused companion projects.

  • Codex Context for ISTM incrementally collects local Codex conversation sessions into ISTM and produces bounded Daily digests with launchd examples on macOS.
  • Mac Context for ISTM collects local Apple Mail, Notification Center, Reminders, and Calendar context into a separate private ISTM store on macOS.
  • Memory Retrieval Lab measures retrieval quality with synthetic fixtures, reproducible metrics, multilingual cases, and ranker adapters.

The reference environment used GPT-5.6 Sol with reasoning effort xhigh on 2026-07-24. The collectors and deterministic baselines are model-independent. Read Daily and ISTM companion projects for the exact repository boundaries and bounded Daily contract.

Provenance and promotion

Promotion is an adjacent-layer, evidence-preserving operation.

06 ISTM -> 05 Daily -> 04 STM -> 03 MTM -> 02 LTM -> 01 XLTM
                         \-------- 00 Life Archive candidate --------/

A promoted record should carry the source pointer, source and capture times, scope, observed fact, derived conclusion, uncertainty, destination, and reason. Promotion does not make a claim true. Mutable facts still require current verification.

The Life Archive arrows above represent selection from structured history, not unrestricted canonical wikilinks. In the current schema, Life Archive links canonically only to adjacent XLTM; nonadjacent evidence remains a plain provenance path.

The reference workflow makes one integrated Structured decision over current XLTM/LTM/MTM/STM plus committed Daily. Parent-before-child is the internal materialization order within that sweep, not a separate parent-first workflow. audit requires every LTM, MTM, and STM record to link to its immediate canonical parent. Same-layer lateral wikilinks are rejected so ownership remains mechanically inspectable.

See Provenance and promotion.

Automation

Memory Forest can be validated and reindexed with POSIX cron, a macOS LaunchAgent, or a Codex Scheduled Task. Reviewed plans can also be applied locally with the same sibling maintenance lock.

Target operating model for automated Memory Forest maintenance

The core does not collect source systems or decide semantic promotions. The automation starter therefore separates two lanes:

  • implemented deterministic maintenance, which locks one exact private root, validates and audits it, then atomically rebuilds the derived index
  • caller-owned integrated layer and structure judgment, followed by the implemented apply-daily or apply-structured transaction writer with provenance binding, rollback, validation, audit, index, and receipt proof

The repository includes a shared maintenance wrapper, a crontab example, a per-user macOS launchd template, and a bounded Codex Scheduled Task prompt. Read the Automation guide before enabling unattended runs. Cron and launchd can remain local. A Codex task uses the selected account and model processing boundary, so do not send strictly local-only evidence through that lane.

Relationship to Codex Debug Bridge

Codex Debug Bridge includes a small memory-forest-starter as an onboarding path for a private route-only helper. This repository is the standalone portable reference implementation. It adds the CLI, contracts, audits, synthetic example, documentation, and companion skill.

Neither project contains a real private forest. Neither project includes private prompts, unattended source collection or semantic-plan generation, operational logs, personal adapters, or user identifiers. The projects can be used independently.

See Codex Debug Bridge integration.

Project status

Memory Forest is alpha software. The v0.3 scope includes strict receipt-backed Daily application, bounded current-Forest context, one integrated XLTM/LTM/MTM/STM sweep, deterministic root-first retrieval, portable contracts, synthetic fixtures, and a route-first body boundary. It is suitable for evaluation and adaptation, not unattended high-stakes decision-making.

Candidate directions include measured multilingual routing evaluation, optional ranking adapters, generated graph views, migration tooling, and stronger provenance integrity checks. These are not release promises. See Roadmap.

Repository map

Path Purpose
src/memory_forest/ local CLI and reference implementation
docs/ architecture, contracts, privacy boundary, and integration guidance
docs/retrieval-guide.md query intake, deterministic retrieval, explicit body boundary, and integration limits
docs/daily-and-istm-companions.md bounded Daily contract and companion repository boundaries
examples/synthetic-forest/ fictional 00-06 forest with no private identifiers
examples/automation/ cron, launchd, and Codex Scheduled Task maintenance examples
examples/memory-forest-retrieve/ metadata-only mandatory consultation gate and prompt
skills/memory-forest/ companion Codex skill
skills/memory-forest-retrieve/ opt-in always-on retrieval integration skill
tests/ behavior and regression tests
scripts/audit_public_release.py public-tree secret and identifier audit

Contributing

Read CONTRIBUTING.md. Keep fixtures synthetic, keep real forests outside Git, and run the full local check before opening a pull request.

make check

Security issues should follow SECURITY.md. Community participation follows CODE_OF_CONDUCT.md.

License

Memory Forest is released under the MIT License.

About

A verifiable local memory architecture for long-running AI agents

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages