A Neo4j knowledge-graph RAG engine for the maritime domain — vessels, operators, ports, regulations and incidents — a ground-truth multi-hop retrieval benchmark (vector 0.70 → best-of ensemble 0.90), and cross-document causal insights from 139 real KMST casualty adjudications
Maritime questions are relational by nature: "Which operators call container ships at Busan?" joins operator→vessel→port; "Which environmental rule applies to the ship that had the Ulsan incident?" chains incident→vessel→type→regulation. Pure vector search retrieves documents about each entity but cannot join them. This project builds a two-layer Neo4j graph (documents + typed entities) over maritime news, serves it through a best-of retriever ensemble (FastAPI + React), and measures how much the graph actually helps.
Real news gives no ground truth, so the repo ships a synthetic maritime world (6 fictional operators, 14 vessels, 6 ports, 4 regulations, 6 incidents) where every relation is known by construction. 42 news-style articles express those relations in prose; 20 QA pairs are derived from the relation tables — including multi-hop questions whose answer appears in no single article. A no-retrieval control proves the fictional entities cannot be answered from GPT-4o's parametric memory.
Same LLM (gpt-4o, temp 0), same answer prompt; only retrieval varies.
Score = gold entities present in the generated answer. Reproduce with
python evaluation/retrieval_benchmark.py:
| Retrieval | Entity recall | Strict acc | 1-hop (n=11) | 2-hop (n=8) | 3-hop (n=1) |
|---|---|---|---|---|---|
| No retrieval (control) | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Vector (top-5 chunks) | 0.73 | 0.70 | 0.91 | 0.38 | 1.00 |
| Graph expansion (VectorCypher) | 0.82 | 0.75 | 0.91 | 0.50 | 1.00 |
| Text2Cypher (direct graph query) | 0.85 | 0.85 | 1.00 | 0.75 | 0.00 |
| LLM router (pick one upfront) | 0.45 | 0.45 | 0.64 | 0.12 | 1.00 |
| Ensemble (run all + judge) | 0.92 | 0.90 | 1.00 | 0.75 | 1.00 |
What the numbers say:
- Vector search collapses on multi-hop (0.91 → 0.38 strict accuracy going from 1-hop to 2-hop): the joined answer set exists in no retrievable chunk.
- Graph expansion recovers part of it (0.50): appending entity facts
(
X -OPERATES- Y) from mentioned entities to each retrieved chunk lets the LLM do some joins in-context. - Direct graph querying wins overall among single retrievers (0.85 strict, 1.00 on 1-hop) — when the question maps to a Cypher pattern, the join happens in the database, not the LLM. But it is brittle on complex chains: the 3-hop question produced valid-looking but empty Cypher. No single retriever dominates.
- Upfront routing makes things worse, not better (0.45 strict — below every single retriever). A router must predict the best retriever before seeing any results; when it guesses wrong, the chosen retriever returns nothing useful and the answer is "정보 없음" (11/20 questions). The prediction is exactly the hard part.
- Selection after retrieval fixes this (0.90 strict, best in every hop bucket): run all three retrievers in parallel, let an LLM judge score each context for evidential support, and generate from the winner. An empty or irrelevant context can never be chosen, so a single retriever's blind spot stops being a failure mode. This is the app's current architecture; the judge picked text2cypher 13×, graph 6×, vector 1× across the 20 questions (avg 3.7s/question vs 1.2–1.9s for single retrievers — the price of running everything, mitigated by parallel execution).
Found along the way (fixed and kept honest): grounding Text2Cypher answers requires
returning the subject alongside the value (RETURN v.name, v.type — a bare
v.type row made the answer LLM refuse), and the auto-extracted Neo4j schema
produced malformed Cypher — the curated schema text in app.py is what the
benchmark validated.
The synthetic benchmark above validates the method; this part applies it to real documents and creates knowledge that did not exist before. The Korean Maritime Safety Tribunal (KMST) publishes an adjudication report for every investigated marine casualty — each narrating a single accident's vessels, location, weather, causal chain of findings, and sanctions. 139 reports (2025–2026, all regional tribunals) were parsed (PDF/HWP/HWPX), extracted with an LLM onto a fixed cause taxonomy, and loaded as a graph:
(Accident {type, night, weather})-[:INVOLVES]->(AVessel {type, tonnage})
(Accident)-[:HAS_CAUSE]->(Cause)-[:OF_TYPE]->(CauseCategory)
(Cause)-[:LEADS_TO]->(Cause) # the adjudicated causal chain
(Accident)-[:IMPOSED]->(Sanction) (Accident)-[:CITES]->(Law)
Every report explains only its own accident. The graph answers questions that exist in no single document:
| Cross-document question | Answer from the graph (139 cases) |
|---|---|
| Which causal chain repeats the most? | "Work-safety-rule violation → inadequate safety-management system": 22 accidents. Individual errors are systematically adjudicated as symptoms of organizational failure. |
| What dominates collision findings? | Lookout negligence: 44 cause findings across 35 collisions — equally frequent by day and night (24 vs 24), while ship-handling and weather causes cluster at night. |
| Fishing vessels vs merchant ships? | Fishing-vessel accidents dominate lookout negligence (38 vs 14) and maintenance neglect (13 vs 2) — a materially different risk profile. |
| Which causes draw the heaviest sanctions? | Cargo-securing failures: average 6.6 months of license suspension, vs 1.8 months for lookout negligence. |
Reproduce: ingest/fetch_kmst_verdicts.py (or drop manually downloaded verdict
files into data/kmst/manual/ and run ingest/parse_manual_verdicts.py) →
ingest/extract_accidents.py → graph/build_accident_graph.py →
evaluation/accident_insights.py. The extracted structured records
(data/kmst/accidents_graph.json) are committed, so the graph and analysis are
reproducible without refetching or an OpenAI key. Source documents are
public-sector works (KOGL) of the Korean Maritime Safety Tribunal. The application (below) runs entirely on this real corpus; the synthetic world
of Part 1 is used only by the retrieval benchmark.
Document layer (Article)-[:HAS_CHUNK]->(Content {embedding})
(Article)-[:PUBLISHED_BY]->(Media) (Article)-[:BELONGS_TO]->(Category)
(Article)-[:MENTIONS]->(entity)
Knowledge layer (Company)-[:OPERATES]->(Vessel {type})
(Vessel)-[:CALLS_AT]->(Port)
(Regulation)-[:APPLIES_TO]->(Vessel)
(Vessel)-[:INVOLVED_IN]->(Incident)-[:OCCURRED_AT]->(Port)
For the bundled corpus the knowledge layer loads from ground-truth tables
(data/corpus/entities.json); for real documents ingest/extract_entities.py
produces the same structure with an LLM extractor.
The application serves the real KMST corpus (Part 2): a FastAPI backend runs three retrievers (verdict-text vector search / graph-expanded context / direct Text2Cypher) in parallel for every question over the 139-adjudication graph; an LLM judge scores each context for evidential support and the answer is generated from the winner, with the judge's choice and scores returned in the response.
Every answer returns its evidence subgraph — the retrieved verdict chunks, the accidents they belong to, and the causes, vessels and locations they connect to — rendered as an interactive force-directed view (d3-force) so the retrieval path is visible, not implied:
| Landing | Single-accident question — chunks → accident → causes |
|---|---|
![]() |
![]() |
# 0. Neo4j
docker run -d --name maritime-neo4j -p 7474:7474 -p 7687:7687 \
-e NEO4J_AUTH=neo4j/maritime123 neo4j:5
# 1. Python env + keys
pip install -r requirements.txt
cp .env.example .env # set OPENAI_API_KEY, NEO4J_PASSWORD
# 2. Generate the synthetic corpus and build the graph (embeds 42 chunks)
python ingest/generate_corpus.py
python graph/build_graph.py
# 3. Run the benchmark
python evaluation/retrieval_benchmark.py
# 4. Serve
uvicorn app:app --port 8001 # backend
cd frontend && npm install && npm run dev # frontend at localhost:5173| Path | Contents |
|---|---|
ingest/generate_corpus.py |
Synthetic maritime corpus + relation tables + QA benchmark (deterministic) |
ingest/extract_entities.py |
LLM entity/relation extractor for real documents (same output schema) |
graph/build_graph.py |
Two-layer Neo4j build: documents, chunks, embeddings, typed entities |
evaluation/retrieval_benchmark.py |
4-way retrieval comparison with gold-entity scoring |
ingest/fetch_kmst_verdicts.py, ingest/parse_manual_verdicts.py |
KMST verdict collection (web + manual PDF/HWP/HWPX parsing) |
ingest/extract_accidents.py |
LLM extraction of causal chains onto a fixed taxonomy |
graph/build_accident_graph.py, evaluation/accident_insights.py |
Accident-layer build + cross-document causal analysis |
app.py |
FastAPI + best-of retriever ensemble (parallel run + LLM judge) + citation-grounded generation |
frontend/ |
React (Vite) search UI |
data/ |
Committed corpus, entity tables, QA benchmark |
- The benchmark corpus is synthetic and small (42 articles, 20 QA); numbers compare retrieval strategies under identical conditions rather than estimate absolute production quality. Articles are templated prose — real news adds paraphrase and noise that would lower all rows, likely vector most.
- Scores are a single run at temperature 0; Text2Cypher failures in particular vary with prompt wording.
- Part 2 relies on LLM extraction (gpt-4o): cause-category assignment is not human-validated, and the corpus is 139 recent cases (2025–26), not a longitudinal sample — read the insights as demonstrations of the graph's analytical reach, not as maritime-safety statistics.
- The knowledge layer for the demo comes from ground truth, isolating retrieval
quality from extraction noise. With
extract_entities.pyon real data, extraction errors compound on top of these numbers.
- Parse-Everything — self-healing document parsing (upstream of any RAG corpus)
- Vehicle-Anomaly-Algorithm · AIS-Traffic-Model · CBM-Anomaly-Dashboard — the maritime/transport AI line this project extends into knowledge retrieval




