Yarr, the docs give up their secrets to no one but us.
A fully local, knowledge-graph-augmented RAG system for HTML, PDF, and plain-text documentation. Runs entirely on Apple M1 Max (64GB). Zero cloud calls. Zero telemetry.
ARG indexes a directory of HTML, PDF, and plain-text documentation, follows links recursively to build a knowledge graph of the corpus, and answers questions over it using hybrid retrieval: dense embeddings (nomic-embed-text via Ollama), sparse BM25 keyword search, and knowledge-graph traversal. Generation uses Qwen3.6 35B (qwen3.6:35b-a3b-q4_K_M) served locally by Ollama. Everything runs on-device — there are no outbound network calls during operation.
Beyond Q&A, ARG provides DCI ("Direct Corpus Interaction") features for
browsing, summarising, and comparing documents in a corpus. The corpus is live:
a watchdog-backed file watcher re-indexes changed files automatically and
updates topic clusters incrementally (only re-labeling the clusters whose
membership changed). ARG exposes both a CLI (scripts/index_docs.py) and a
local web UI (FastAPI bound to 127.0.0.1:8000), and supports multiple
independent corpora via a path-convention namespace (arg_db/{corpus_name}/).
Retrieval is a 5-stage pipeline:
- Stage 0 — DCI enrichment. BM25 chunk scores aggregated per document find the most relevant files; their link neighbourhoods and topic clusters expand the candidate pool.
- Stage 1 — Dense retrieval. Top-k chunks via nomic-embed-text + ChromaDB.
- Stage 1.5 — BM25 retrieval. Sparse exact-term recall via bm25s (Rust-backed), run in parallel with dense.
- Stage 2 — Graph traversal. Kuzu-backed link expansion: neighbours of the dense/sparse hits are added as additional candidates.
- Stage 3 — RRF fusion. Reciprocal Rank Fusion (k=60) merges all ranked lists.
- Stage 4 — Lost-in-middle reordering. Top-ranked chunks moved to the head and tail of the prompt; mid-ranked chunks placed in the middle.
The fused, reordered context is handed to Qwen3.6 35B via LlamaIndex's query engine.
- Apple M1 Max (or any Apple Silicon with sufficient RAM; 32GB minimum)
- macOS 13+
- Homebrew
- Python 3.11+
- ~25GB free disk space (~22GB for the qwen3.6:35b-a3b-q4_K_M model, rest for index data)
bash scripts/bootstrap.shBootstrap checks for Ollama and installs it via Homebrew if missing; pulls
qwen3.6:35b-a3b-q4_K_M and nomic-embed-text only if they are not
already in ollama list; installs Tesseract (for pymupdf OCR tessdata);
downloads D3.js v7 once to arg/static/d3.min.js; creates a Python 3.11+
virtual environment at .venv/ using stdlib python3.11 -m venv; activates it;
and installs the project via pip install -e .[dev]. If a ./vendor/ wheel
cache is present, pip uses it offline (--no-index --find-links ./vendor/);
otherwise it falls back to PyPI. Bootstrap is the only step that touches the
network — once it completes, ARG runs entirely offline.
source .venv/bin/activate# 1. Index a documentation directory
python scripts/index_docs.py index --docs /path/to/your/docs --db ./arg_db
# 2. Start the web UI
python scripts/index_docs.py serve --db ./arg_db
# 3. Open http://localhost:8000 in your browserpython scripts/index_docs.py index --docs PATH --db PATH [--corpus NAME] [--no-watch]
[--subset SUBDIR] [--include PATTERN] [--reset]
python scripts/index_docs.py query --db PATH [--corpus NAME] [--no-enrich]
python scripts/index_docs.py serve --db PATH [--corpus NAME] [--port PORT]
python scripts/index_docs.py stats --db PATH [--corpus NAME]Index only a subdirectory or file type without waiting for the full corpus:
# Re-index just one folder (wipe first)
python scripts/index_docs.py index --docs ~/index --db ./index_db \
--subset ~/index/Retirement --reset --no-watch
# Re-index only PDFs
python scripts/index_docs.py index --docs ~/index --db ./index_db \
--include "*.pdf" --reset --no-watch
# --include is repeatable; multiple patterns use OR logic
python scripts/index_docs.py index --docs ~/index --db ./index_db \
--subset ~/index/Retirement --include "*.html" --include "*.pdf" --no-watch--reset deletes the corpus before indexing (no confirmation prompt when the flag is
explicit). Without --reset the filtered files are merged into the existing index.
python scripts/index_docs.py index --corpus product_a --docs /path/to/product_a --db ./arg_db
python scripts/index_docs.py index --corpus product_b --docs /path/to/product_b --db ./arg_db
python scripts/index_docs.py serve --db ./arg_db # ?corpus= param selects corpus per request# Interactive reset (asks you to type the corpus name to confirm)
python scripts/reset_corpus.py --db ./arg_db --corpus default
# Non-interactive reset (skip prompt — same as passing --reset to index)
python scripts/reset_corpus.py --db ./arg_db --corpus default --confirmCreate eval/qa_pairs.json with hand-written question/answer pairs, then:
python scripts/eval_retrieval.py --db ./arg_db --corpus default --qa eval/qa_pairs.jsonOpen http://localhost:8000 after starting the server. The UI provides:
- Ask — type a question; sources are shown as clickable links that open the document in a new tab
- Topic clusters — visual map of documents grouped by topic (scroll to zoom, drag to pan, click a shape to open the file). Shapes indicate file type: circle = PDF, square = HTML, triangle = text/markdown, diamond = other. Clusters are computed at the end of
indexand kept current by the watcher. - Documents — sidebar list of all indexed files; click to see key points and metadata
- Topics — sidebar list of LLM-generated cluster labels
All tunables live in .env (copy from .env.example). Key settings:
CHUNK_SIZE=1024— token window per chunkENRICH_MIN_SCORE=0.5— DCI enrichment threshold (tune witheval_retrieval.py)QUERY_REWRITE=false— rewrite conversational queries (disable for speed)QUERY_DECOMPOSE=false— decompose compound queries (disable for speed)BM25_ENABLED=true— sparse exact-term retrieval alongside denseWATCH_ENABLED=true— auto re-index files changed on diskOLLAMA_TIMEOUT=300— seconds before an Ollama call times out
Logs are written to {db}/{corpus}/arg.log in JSON format. Key messages to watch for:
| Logger | Message | Meaning |
|---|---|---|
arg.pipeline |
pipeline: ready (…query_rewrite=… query_decompose=…) |
Startup — confirms env-var overrides took effect |
arg.pipeline |
watcher: modified foo.pdf |
File change detected; re-indexing triggered |
arg.retriever |
retriever: BM25 index reloaded |
BM25 updated after watcher re-index |
arg.generator |
generator: query answered in 4321ms — "…" |
Query completed |
arg.dci.explorer |
explorer: labeling cluster c2 (16 docs) via LLM |
Cluster label being generated |
arg.dci.analyst |
analyst: summary cache hit for foo.pdf |
Summary served from cache (no LLM call) |
Stream logs live with tail -f {db}/default/arg.log. Add --debug to the serve command for verbose tracing.
After indexing, generate a pair of self-contained HTML analysis reports from the log files:
python scripts/index_report.pyBy default this reads ./index_db/default/ and writes to ./reports/. Both paths are configurable:
python scripts/index_report.py --db ./index_db --corpus default --out ./reports
# Open the main report in your browser when done
python scripts/index_report.py --openTwo files are produced:
| File | Contents |
|---|---|
reports/indexing_report.html |
Performance metrics, error breakdown, extraction quality, cluster summary, and prioritised improvement recommendations |
reports/indexing_appendix.html |
Full raw-data tables: every encrypted PDF skipped, every 0-chunk document, every AcroForm warning, slowest 50 documents, all indexing gaps, hourly throughput |
The report covers:
- Performance — wall time vs. CPU time, per-document latency percentiles (p50/p95/p99), slowest documents, OCR usage
- Errors & warnings — distinction between benign library noise (pdfminer font warnings) and real application warnings (encrypted PDFs, dangling links, AcroForm partial extraction)
- Extraction quality — documents that produced 0 chunks (unreachable in retrieval), 1-chunk documents, very large documents that dominate the index
- Clusters — LLM-generated topic labels and document counts
- Index artifacts — ChromaDB embedding counts, Kuzu graph node/edge counts, BM25 corpus size
The reports/ directory is git-ignored; reports are re-generated on demand from the live log files.
-
Only four file types are indexed —
.pdf,.html/.htm,.txt, and.md. Everything else in the corpus directory is silently skipped. This includes:- Images (
.jpg,.jpeg,.png,.heic,.tiff,.gif,.webp,.svg) - Office documents (
.docx,.xlsx,.pptx,.doc,.xls) - Stylesheets, scripts, and web fonts (
.css,.js,.woff,.woff2) - Resource files inside browser-saved-page directories (e.g.
PageName_files/) - Archives (
.zip,.tar,.gz) - Audio and video files (
.mp3,.mp4,.mov,.m4a) - macOS metadata files (
.DS_Store,._*) - Any file in a hidden directory (names starting with
.)
As a result,
find . -type f | wc -lon your corpus will report a larger number than ARG's indexed document count. To get a comparable count:find . -type f \( -iname "*.pdf" -o -iname "*.html" -o -iname "*.htm" \ -o -iname "*.txt" -o -iname "*.md" \) | wc -l
- Images (
-
Markdown structure is not parsed —
.mdand.markdownfiles index as plain text. Atx-style headings (# H1,## H2) are not recognised as chunk boundaries. A future feature may add Markdown-aware extraction. -
iframes are not followed or indexed — content inside an
<iframe src=...>is invisible to the crawler. -
JavaScript-rendered content is not indexed — ARG parses the HTML as served; it does not execute scripts.
-
Encrypted PDFs are skipped (logged with a clear reason); ARG does not prompt for passwords.
-
Right-to-left scripts (Arabic, Hebrew) are not specifically supported — text extracts but reading order may be wrong for mixed-direction layouts.
-
Scanned-PDF accuracy depends on Tesseract's
engtessdata; non-English scanned PDFs need the matchingtessdatapack installed via Homebrew. -
External links (
http://,https://) are recorded as graph edges but never fetched — onlyfile://and relative paths are crawled. -
Authentication / access control is out of scope; unreadable files are logged and skipped rather than retried.
pytest tests/unit/ # fast; no Ollama required (LLM mocked)
pytest tests/integration/ # requires indexed corpus; LLM mocked
pytest tests/e2e/ # requires running Ollama with qwen3.6 and nomic-embed-text