-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathenv.example
More file actions
421 lines (393 loc) · 19.3 KB
/
Copy pathenv.example
File metadata and controls
421 lines (393 loc) · 19.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
### MemGraphRAG environment sample — the single template for this repository.
### Copy to .env and edit: `cp env.example .env`.
### OS environment variables take precedence over .env.
###
### Project-owned vars use the MEMGRAPHRAG_ prefix; infrastructure/binding vars
### keep LightRAG naming (POSTGRES_*, NEO4J_*, LLM_*, EMBEDDING_*, DOCLING_*, CHUNK_*).
###
### Every variable below is read by the code. Settings that were documented but
### never implemented were removed; the few that are kept for an external service
### (not for MemGraphRAG itself) say so explicitly. Verify with:
### grep -rn "NAME" memgraphrag --include='*.py'
###
### The server does NOT read .env on import: entry points call
### `memgraphrag.api.config.load_env_file()`. A .env sitting next to a process that
### merely imports the package (pytest, a script) is therefore ignored.
###########################################################################
### Reading this file from Docker
###########################################################################
### Both compose files load this whole file with `env_file:`, so every setting
### below reaches the container. What they do NOT let through is anything that
### names *this machine*, because `environment:` wins over `env_file:`:
###
### docker-compose.yml self-contained stack. Also overrides POSTGRES_HOST,
### POSTGRES_PORT and NEO4J_URI to its own services —
### pointing them at a LAN host here has no effect.
### docker-compose.app.yml app + web UI against databases you already run.
### POSTGRES_* / NEO4J_* / WORKSPACE come from here.
###
### Both always override: HOST, PORT, WORKING_DIR, INPUT_DIR, LIBRARY_ROOT and the
### four certificate paths — they are host-relative here and meaningless in /app.
### See docs/DockerDeployment.md.
###########################
### Server Configuration
###########################
HOST=0.0.0.0
PORT=9621
# Docker image tag for compose builds (exeio-memgraphrag:<version>)
# MEMGRAPHRAG_VERSION=0.1.0
# WORKERS: gunicorn worker processes. Startup FAILS with WORKERS>1 while any
# file-backed backend (JsonKVStorage / NanoVectorDBStorage / IgraphStorage /
# JsonDocStatusStorage) is selected: their locks are asyncio locks inside one
# process, so two workers sharing WORKING_DIR corrupt the JSON/GraphML files.
# Even on Postgres/Neo4j the ingest lock stays per-process, so "409 while busy"
# only holds within one worker. Leave this at 1 unless you know why.
# WORKERS=1
# CORS_ORIGINS=*
# CORS_ORIGINS="*" disables credentialed cross-origin requests (a reflected
# Origin + credentials would let any page drive this API). Name explicit origins
# to allow cookies/Authorization from a browser.
WORKING_DIR=./data/rag_storage
INPUT_DIR=./data/inputs
# LOG_LEVEL=INFO
# Seconds the shutdown handler waits for in-flight indexing to finish before it
# cancels the workers. Bounded on purpose: an unbounded wait turns a rolling
# restart into a hang, and the container runtime SIGKILLs anyway.
# MEMGRAPHRAG_SHUTDOWN_DRAIN_TIMEOUT=30
# SSL=false
# SSL_CERTFILE=
# SSL_KEYFILE=
# Outbound TLS (LLM/embed/Docling/clients URL-ingest). Corporate inspection:
# Place your Fortinet/Zscaler CA at ./certs/corporate-ca.crt (gitignored), or set:
# MEMGRAPHRAG_SSL_CERT_FILE=/path/to/Fortinet_CA_SSL.crt
# SSL_CERT_FILE=/path/to/Fortinet_CA_SSL.crt
# SSL_VERIFY=true
# SSL_VERIFY=false # lab-only escape hatch (insecure)
# OpenSSL 3 / Python 3.13: custom CA paths auto-relax VERIFY_X509_STRICT (AKI).
###########################
### Auth (JWT and/or API key)
###########################
# AUTH_ACCOUNTS=admin:change_me
# TOKEN_SECRET is mandatory as soon as AUTH_ACCOUNTS is set: the fallback secret
# is published in this repository, so tokens signed with it are forgeable.
# TOKEN_SECRET=change-me-to-a-long-random-string
# JWT_ALGORITHM=HS256 ("none" is rejected)
# TOKEN_EXPIRE_HOURS=48
# GUEST_TOKEN_EXPIRE_HOURS=24
# MEMGRAPHRAG_API_KEY=
# WHITELIST_PATHS=/health,/docs,/openapi.json
# WARNING: never add /api/* here. The Ollama router is mounted on /api and its
# /api/chat and /api/generate routes call the billed LLM (/bypass skips retrieval
# entirely). The whitelist short-circuits BEFORE any token or API-key check.
# Fail closed: refuse to serve unauthenticated requests even if the .env failed to
# load (e.g. the process was started from another working directory).
# REQUIRE_AUTH=false
# POST /login rate limit, per client IP (429 past the cap).
# LOGIN_MAX_ATTEMPTS=10
# LOGIN_WINDOW_SECONDS=60
# Hard cap on a single uploaded document, in bytes (default 100 MiB → 413).
# MAX_UPLOAD_SIZE=104857600
###########################
### Query / retrieval (MemGraphRAG-native)
###########################
TOP_K=10
LINKING_TOP_K=50
PASSAGE_NODE_WEIGHT=0.05
DAMPING=0.5
FACT_SIMILARITY_THRESHOLD=0.6
SKIP_FACT_RERANK=true
SCHEMA_TOP_K=5
SCHEMA_NODE_WEIGHT=0.1
PPR_ENGINE=igraph
###########################
### Ontology / conflict construction (index-time)
###########################
ONTOLOGY_BATCH_SIZE=20
# Chunks extracted between two OpenIE cache writes. A run killed mid-way keeps every
# completed sub-batch; the relaunch re-bills only the sub-batch that was in flight.
OPENIE_CHECKPOINT_SIZE=64
ONTOLOGY_MIN_FREQUENCY=2
# Safety valve for the frequency filter. If pruning schemas below the minimum
# frequency would deactivate more than this share of the facts, the filter is
# skipped for that build and a warning names the ratio — a fragmented schema layer
# (mixed languages, unfolded accents) must not silently empty the fact graph.
ONTOLOGY_MAX_DEACTIVATION_RATIO=0.5
# Language of extracted entities, relations and type labels, and of answers.
# "auto" leaves it to the model, which on a non-English corpus mixes languages:
# it emits ("Entreprise","doit emettre","facture") next to ("Company","must
# issue","invoice"), so every concept gets two schemas at frequency 1 and the
# ontology filter deactivates itself. Set to the corpus language, e.g. "French".
# MEMGRAPHRAG_LANGUAGE=auto
CONFLICT_ENABLED=true
CONFLICT_MAX_GROUPS=50
# Minimum judge confidence below which a conflict resolution is discarded.
CONFLICT_MIN_CONFIDENCE=0.85
# Characters of conflicting evidence handed to the resolver in one prompt
# (floored at 2000 by the code).
CONFLICT_RESOLUTION_CHAR_BUDGET=24000
###########################
### LLM binding (OpenAI-compatible)
###########################
### Default: public OpenAI API. Any OpenAI-compatible gateway works
### (Azure, vLLM, Ollama's /v1 shim, …) — point LLM_BINDING_HOST at it.
LLM_BINDING=openai
LLM_BINDING_HOST=https://api.openai.com/v1
LLM_BINDING_API_KEY=
LLM_MODEL=gpt-4o-mini
MAX_ASYNC_LLM=4
# Comma-separated allow-list for the per-request `model` field on /query and the
# picker in the web UI. LLM_MODEL is always offered even when this is empty; any
# other value is rejected with 400 rather than forwarded to the provider. Only list
# models your LLM_BINDING_HOST actually serves.
# LLM_MODELS=gpt-4o-mini,gpt-4o
### Per-request provider routing (completions only)
# The web UI can route a single query to another OpenAI-compatible endpoint. A
# request names a provider *id*; the server resolves the credential from the
# variables below, so a browser never carries a key. Read by memgraphrag/llm/providers.py.
# Each provider takes an optional <ID>_MODELS allow-list; when it is empty the
# endpoint's own GET /v1/models catalogue is used instead.
#
# EMBEDDINGS ARE NEVER ROUTED. The corpus is indexed with one embedding model at
# one dimension; sending query embeddings elsewhere returns vectors from a
# different space and silently degrades every answer.
# TOGETHER_API_KEY=
# TOGETHER_MODELS=openai/gpt-oss-20b,meta-llama/Llama-3.3-70B-Instruct-Turbo
# TOGETHER_BASE_URL=https://api.together.ai/v1
# OLLAMA_BASE_URL=http://localhost:11434/v1
# OLLAMA_MODELS=llama3.2,qwen2.5
# VLLM_BASE_URL=
# VLLM_MODELS=
### Document library browsed by the web UI (read-only)
# Root folder served by GET /library/tree, /library/file and /library/preview.
# Every requested path is resolved and must stay under this root.
# LIBRARY_ROOT=./data/library
# Under Docker, LIBRARY_ROOT is forced to /app/data/library and LIBRARY_HOST_DIR is
# what gets mounted there, read-only. Setting LIBRARY_ROOT for a compose run has no
# effect; set this instead.
# LIBRARY_HOST_DIR=./data/library
### MCP server (optional, read-only, mounted on the API's own port at /mcp)
# Exposes retrieve / search_documents / read_document / cypher to third-party MCP
# clients. Off by default; see docs/MCP.md.
# MCP_ENABLED=false
# Exact-match host allow-list for the transport's DNS-rebinding protection. Leave
# empty for localhost only. A remote deployment MUST list its hostname, with and
# without the port, or every call comes back 421 Invalid Host header.
# MCP_ALLOWED_HOSTS=rag.example.com,rag.example.com:9621
# Advertised as the issuer/resource URL in the OAuth metadata the SDK serves.
# MCP_ISSUER_URL=http://localhost:9621
### Agent mode (mode=agent on /query and /query/stream)
# Token ceiling for the loop's message list before old tool results are evicted.
# AGENT_CONTEXT_BUDGET=24000
# Output ceiling for a *deciding* turn — the call that only chooses whether to
# search again. Its prose is discarded, so an uncapped one is pure latency: measured
# at 37 s and 4 606 characters of unread reasoning on the reference corpus, against
# 2.7 s at this default. Raise it only if a model needs room to reason before
# emitting a tool call; the server warns when a turn is cut off without having made
# one. The answering turn is never capped.
# AGENT_DECIDE_MAX_TOKENS=256
# Comma-separated substrings of models allowed to drive the loop. When set it is an
# allow-list; when empty, only models that plainly cannot call tools are refused.
# Whether a model performs a *second* search is a property of the model, not of the
# code: Llama-3.3-70B-Instruct-Turbo and Kimi-K3 re-search with a reformulated query,
# while neither gpt-oss variant ever does. See docs/WebUI.md for the measurements.
# AGENT_TOOL_MODELS=Llama-3.3-70B,Kimi-K3
# Extra substrings to refuse, on top of the built-in ones.
# AGENT_TOOL_DENY=
# OPENAI_API_KEY=
### NOT CONFIGURABLE: the LLM HTTP timeout is hard-coded in
### memgraphrag/llm/openai_compatible.py as httpx.Timeout(150.0, connect=30.0).
### An LLM_TIMEOUT variable is read nowhere; slow self-hosted models need a code
### change, not an env var.
### Local Ollama (OpenAI-compatible API, typical port 11434):
# LLM_BINDING_HOST=http://localhost:11434/v1
# LLM_BINDING_API_KEY=ollama
# LLM_MODEL=llama3.2
### Self-hosted vLLM (e.g. Mistral GPTQ INT4, --served-model-name mistral):
# LLM_BINDING_HOST=http://localhost:8001/v1
# LLM_BINDING_API_KEY=EMPTY
# LLM_MODEL=mistral
# MAX_ASYNC_LLM=8
### These are the only variables that reach MemGraphRAG. VLLM_* / HF_TOKEN style
### settings belong to your vLLM container's own compose file — this repository
### ships no vLLM service and reads none of them.
###########################
### Embedding binding (OpenAI-compatible)
###########################
EMBEDDING_BINDING=openai
EMBEDDING_BINDING_HOST=https://api.openai.com/v1
EMBEDDING_BINDING_API_KEY=
EMBEDDING_MODEL=text-embedding-3-small
EMBEDDING_DIM=1536
# Cap embed inputs (provider limit). Tiktoken undercounts e5/bge; safety shrinks budget.
# EMBEDDING_MAX_TOKENS=512
# Request-level batching. EMBEDDING_MAX_TOKENS bounds each *text*; these bound the
# *request*, which nothing did before: a corpus-sized call (1 700 chunks x 1 200
# tokens) overflowed any provider ceiling and failed the whole document. A refused
# batch is halved and retried, so these are safety rails, not tuning knobs.
# EMBEDDING_BATCH_SIZE=64
# EMBEDDING_BATCH_MAX_TOKENS=100000
# EMBEDDING_TOKEN_SAFETY=0.60
# EMBEDDING_SEND_DIMENSIONS=true
### Local Ollama embeddings example:
# EMBEDDING_BINDING_HOST=http://localhost:11434/v1
# EMBEDDING_BINDING_API_KEY=ollama
# EMBEDDING_MODEL=nomic-embed-text
# EMBEDDING_DIM=768
### Note: a chat-only vLLM service cannot serve embeddings — keep a separate
### embedding endpoint (OpenAI / Together / Ollama / an embedding vLLM).
### NOT IMPLEMENTED: EMBEDDING_QUERY_PREFIX. Asymmetric query prefixes come from
### the linking prompts in memgraphrag/prompts/, not from an env var.
###########################
### Storage selection
###########################
### Defaults (no external DB, single worker only):
### JsonKVStorage, NanoVectorDBStorage, IgraphStorage, JsonDocStatusStorage
MEMGRAPHRAG_KV_STORAGE=JsonKVStorage
MEMGRAPHRAG_VECTOR_STORAGE=NanoVectorDBStorage
MEMGRAPHRAG_GRAPH_STORAGE=IgraphStorage
MEMGRAPHRAG_DOC_STATUS_STORAGE=JsonDocStatusStorage
# WORKSPACE=
# Move a JSON KV file aside instead of failing when it is unreadable.
# MEMGRAPHRAG_KV_QUARANTINE_CORRUPT=false
### Compose / integration overrides (docker-compose.yml sets these itself):
# MEMGRAPHRAG_KV_STORAGE=PGKVStorage
# MEMGRAPHRAG_VECTOR_STORAGE=PGVectorStorage
# MEMGRAPHRAG_GRAPH_STORAGE=Neo4JStorage
# MEMGRAPHRAG_DOC_STATUS_STORAGE=PGDocStatusStorage
# PPR_ENGINE=neo4j_gds
###########################
### PostgreSQL (+ pgvector)
###########################
POSTGRES_HOST=localhost
POSTGRES_PORT=5432
POSTGRES_USER=rag
# REQUIRED by docker-compose.yml (no default any more).
POSTGRES_PASSWORD=change-me
POSTGRES_DATABASE=rag
# POSTGRES_MAX_CONNECTIONS=10
# POSTGRES_VECTOR_INDEX_TYPE=hnsw
# POSTGRES_HNSW_M=16
# POSTGRES_HNSW_EF=64
# POSTGRES_IVFFLAT_LISTS=100
### Compose publish port (host → container):
# Published on 127.0.0.1 by default; set the ADDR to 0.0.0.0 to expose deliberately.
# POSTGRES_PUBLISH_ADDR=127.0.0.1
# POSTGRES_PUBLISH_PORT=5432
###################################################
### Application data (chat threads) — SEPARATE from
### the RAG's PostgreSQL above. Different lifecycle,
### different backups, different blast radius.
###################################################
# Read by memgraphrag/api/config.py; consumed by memgraphrag/chat/store.py.
# docker-compose.yml runs this as its own `postgres-app` service, published on
# 127.0.0.1:5433 because 5432 is commonly already taken by the RAG database.
# Unset means chat persistence is OFF: /chat/* answers 503 and the web UI keeps
# conversations in the browser tab only. Every other route is unaffected.
# APP_DATABASE_URL=postgresql://app:change-me@127.0.0.1:5433/memgraphrag_app
# REQUIRED by docker-compose.yml for the postgres-app service (no default).
# APP_POSTGRES_PASSWORD=change-me
# REQUIRED by both compose files. Guarded with `:?`, so a missing value stops the
# stack before anything starts rather than leaving /chat/* answering 503.
### Compose publish port (host → container):
# APP_POSTGRES_PUBLISH_ADDR=127.0.0.1
# APP_POSTGRES_PUBLISH_PORT=5433
###########################
### Neo4j (+ GDS plugin for PPR_ENGINE=neo4j_gds)
###########################
NEO4J_URI=neo4j://localhost:7687
# Neo4j labels nodes with the workspace name, and so does LightRAG. Startup refuses
# a workspace already holding nodes MemGraphRAG did not create, because sharing one
# mixes two knowledge graphs in every traversal. clear() is scoped to owned nodes, so
# nothing is destroyed either way. Set to true only to share a workspace knowingly.
# MEMGRAPHRAG_ALLOW_SHARED_NEO4J_WORKSPACE=false
NEO4J_USERNAME=neo4j
# REQUIRED by docker-compose.yml (no default any more).
NEO4J_PASSWORD=change-me
# Published on 127.0.0.1 by default (gds.* runs unrestricted on this instance).
# NEO4J_PUBLISH_ADDR=127.0.0.1
# NEO4J_DATABASE=neo4j
# NEO4J_WORKSPACE=
# NEO4J_MAX_CONNECTION_POOL_SIZE=50
# NEO4J_CONNECTION_TIMEOUT=30.0
### Compose publish ports:
# NEO4J_HTTP_PORT=7474
# NEO4J_BOLT_PORT=7687
###########################
### File processing / chunking
###########################
# MEMGRAPHRAG_PARSER=*:legacy-F
### CHUNK_SIZE / CHUNK_OVERLAP_SIZE apply to every chunker (F, R and P alike):
### memgraphrag/chunker/ reads no environment variable of its own, the pipeline
### passes these two down. Per-chunker CHUNK_F_SIZE / CHUNK_R_SIZE / CHUNK_P_SIZE
### (and their *_OVERLAP_SIZE) were advertised but never implemented — removed.
CHUNK_SIZE=1200
CHUNK_OVERLAP_SIZE=100
# PDF_DECRYPT_PASSWORD=
# TIKTOKEN_MODEL=gpt-4o-mini
### NOT IMPLEMENTED: MAX_PARALLEL_INSERT — ingestion concurrency is fixed in
### memgraphrag/pipeline.py; only MAX_ASYNC_LLM bounds outbound LLM concurrency.
###########################
### Docling (optional remote parser)
###########################
# Local compose profile: docker compose --profile docling up -d
# DOCLING_ENDPOINT=http://localhost:5001
# DOCLING_ADDITIONAL_SUFFIXES=
# DOCLING_POLL_INTERVAL_SECONDS=2
# DOCLING_MAX_POLLS=300
# DOCLING_PUBLISH_ADDR=127.0.0.1
# DOCLING_PUBLISH_PORT=5001
# From inside compose (memgraphrag → docling profile service):
# DOCLING_ENDPOINT=http://docling:5001
### NOT IMPLEMENTED: DOCLING_DO_OCR, DOCLING_FORCE_OCR, MAX_PARALLEL_PARSE_DOCLING,
### MEMGRAPHRAG_FORCE_REPARSE_DOCLING. The client in
### parser/external/docling/parser.py reads only the three DOCLING_* settings above
### and sends no conversion options at all, so OCR follows whatever the
### docling-serve instance itself defaults to.
###########################
### Langfuse observability (retrieval traces)
###########################
# When enabled, query/retrieval stages emit nested Langfuse observations.
# Requires `langfuse` (included in the [api] extra).
# LANGFUSE_ENABLE_TRACE=false
# LANGFUSE_PUBLIC_KEY=
# LANGFUSE_SECRET_KEY=
# LANGFUSE_HOST=https://cloud.langfuse.com
# LANGFUSE_BASE_URL=https://cloud.langfuse.com
###########################
### Ollama emulation (server exposes Ollama-compatible /api/*)
###########################
# OLLAMA_EMULATING_MODEL_NAME=memgraphrag
# OLLAMA_EMULATING_MODEL_TAG=latest
###########################
### Gunicorn (optional; memgraphrag-gunicorn entry point)
###########################
# MEMGRAPHRAG_GUNICORN_BIND=0.0.0.0:9621
# MEMGRAPHRAG_GUNICORN_WORKERS=1 (same WORKERS>1 refusal as above)
# MEMGRAPHRAG_GUNICORN_LOGLEVEL=info
# ACCESS_LOG=-
# ERROR_LOG=-
###########################
### Clients (CLI / Streamlit, run outside the service image)
###########################
# MEMGRAPHRAG_SERVER_URL=http://localhost:9621
###########################
### Removed settings (migrating an older .env)
###########################
### These names used to appear in the templates but are read nowhere in
### memgraphrag/. Leaving them in a .env is harmless but changes nothing:
### ENABLE_LLM_CACHE - there is no LLM cache. NameSpace.KV_LLM_CACHE is
### declared but no storage is ever instantiated on it.
### TIMEOUT, LLM_TIMEOUT - the LLM timeout is hard-coded (see above).
### CHUNK_F_SIZE / CHUNK_R_SIZE / CHUNK_P_SIZE (+ *_OVERLAP_SIZE)
### - only CHUNK_SIZE / CHUNK_OVERLAP_SIZE exist.
### MAX_PARALLEL_INSERT - ingest concurrency is fixed in pipeline.py.
### DOCLING_DO_OCR, DOCLING_FORCE_OCR, MAX_PARALLEL_PARSE_DOCLING,
### MEMGRAPHRAG_FORCE_REPARSE_DOCLING
### - the Docling adapter reads only DOCLING_ENDPOINT,
### DOCLING_POLL_INTERVAL_SECONDS, DOCLING_MAX_POLLS.
### EMBEDDING_QUERY_PREFIX - query prefixes come from memgraphrag/prompts/.
### VLLM_* (13 names), HF_TOKEN, MISTRAL_MODEL_7B_GPTQ_4
### - settings for a vLLM container this repo does not
### ship; MemGraphRAG only reads LLM_* / EMBEDDING_*.
### The old second template `.env.example` was deleted: two divergent samples meant
### `cp env.example .env` silently dropped the ontology/conflict knobs.