Retrieve wuxia source text with RAG, hand it to a multi-agent pipeline that writes a structured wuxia RPG script.
BiXiaScribe is a wuxia (武俠) RPG script generator. Give it one line of a story requirement (e.g. "a Shaolin disciple leaves the temple to investigate a massacre"), and it retrieves relevant passages from your own corpus of wuxia novels, then hands them to three specialized LLM agents (writer → dialogue → proofreader) that produce a structured script JSON — NPCs, events, branching choices, triggers — usable as source material for downstream game production (e.g. RPG Maker).
A browser UI for reading and comparing scripts already generated by scripts/eval_generation.py,
instead of opening out/eval/*.json by hand. Four modes — single-script / side-by-side /
overview table / generate. The generate mode
triggers a real run from the browser on a background thread, with a live elapsed-time clock, a
task-boundary progress bar, and a cancel button that actually interrupts the run.
Events tab: triggers, dialogue lines, and branching choices rendered as readable prose instead of
raw JSON.
Surfacing retrieval_calls per script in the UI is what makes the zero-retrieval-call finding from
Key results below checkable at a glance, instead of something you have to infer from log
scrollback.
- RAG indexing: Chinese-aware chunking → embedding → Chroma, resumable.
- Hybrid retrieval (vector + hand-rolled BM25, see Key results below).
- Layered/stateful script generation pipeline (extract → beats → scene-by-scene → proofread, real-time causal-graph validation, checkpointed resume, batch confirmation), structured JSON + cross-reference validation.
- Per-agent model-split A/B and cost accounting (
scripts/eval_generation.py); unit tests never hit a real API. - Retrieval can be switched off entirely (
RETRIEVAL_ENABLED=false/--no-retrieval/ a UI checkbox) to save the single largest token cost, for A/B'ing whether a model with a strong native wuxia voice actually needs corpus grounding. - Streamlit UI, four modes including browser-triggered generation — see UI preview above. Can delete/export/import scripts, restricts model selection to a curated catalog, sets a global reasoning-effort knob, and resumes an interrupted layered run (hard-gated on checkpoint schema version).
- 📋 Planned: editing script content and saving it back from the UI.
Compared to just prompting ChatGPT directly for a script, BiXiaScribe differs in:
- RAG retrieval over real source text, not just the model's imagination of "wuxia voice" — it indexes your own corpus and feeds retrieved passages into the dialogue agent's prompt, so wording and move names stay closer to the source material.
- A Chinese-aware chunker — measures length in characters and prefers splitting at paragraph/punctuation boundaries, instead of reusing token-splitting logic built for English NLP.
- Structured output with automated cross-reference validation —
dialogue.npc/choices[].nextcross-references are re-checked in Python, not just trusted because the proofreader agent says it's fine. - Local-first, runnable end-to-end at zero cost — the default embedding backend is the
local
bge-m3model (offline, free, no API key); the LLM also has afakemode so tests never make a real API call.
- Retrieval: under a strict comparison (
--top-k 1), term-match rate is 91.7% for hybrid vs. 75% for vector-only — character-bigram BM25 catches proper-noun matches vector search tends to miss. - Generation no-RAG A/B (2026-08-19): the retrieval-on variant's
usd_per_event($0.0043) is lower than the retrieval-off variant's ($0.0081) — despite costing 40% more per run — while also having a higher NPC-speaking rate and longer dialogue lines, and cutting dead-end self-loop branches from 20% to 0%. - A still-valid non-obvious finding:
retrieval_callsshows that "the model supports function calling" per OpenRouter's metadata is not the same guarantee as "it reliably chooses to call the tool in a CrewAI ReAct loop" — this needs checking per model.
Full tables, methodology, and past A/B rounds:
docs/BENCHMARKS.md (in Chinese).
python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # only needed to run real script generation (OpenRouter)# 1. Build an index (sample corpus, done in seconds, no API key)
python scripts/build_index.py --corpus tests/sample_corpus.txt
# 2. Query the index (default: hybrid = vector + BM25)
python scripts/test_retrieval.py --query "獨孤九劍的劍法精要" --top-k 3
# 3. Check the backend/API key/index are wired up before spending a token
python scripts/generate_script.py --requirement "test" --preflight-only
# 4. Generate a script (needs LLM_BACKEND=openrouter + OPENROUTER_API_KEY) --
# checkpointed under .bixia_state/<run_id>/, resumable (see CLAUDE.md)
python scripts/generate_script.py --requirement "少林弟子下山查一樁滅門案" --out script.json
python scripts/generate_script.py --requirement "..." --run-id <run_id> # resume a specific run
# 5. Browse/compare already-generated scripts (no API key, no tokens spent),
# or use the 生成 (generate) mode to trigger a real run from the browser
pip install -r requirements-ui.txt
.venv/bin/streamlit run ui/app.py(screenshots above, see UI preview)
Drop your own .txt files into data/corpus/ (not assumed UTF-8). Full commands for switching
corpora/embedding backends and comparing retrieval/model-split quality are in
docs/DESIGN_NOTES.md (in Chinese).
script.json looks roughly like this (full field definitions in
src/bixiascribe/schema.py):
{
"meta": { "title": "...", "theme": "...", "goal": "...", "tone": "..." },
"stat": { "id": "mood", "name": "mood value", "init": 50 },
"player": { "name": "...", "origin": "...", "flaw": "...", "token": "..." },
"items": [{ "id": "...", "name": "...", "from_event": "..." }],
"npcs": [{
"id": "...", "name": "...", "faction_id": "...", "role": "...",
"personality": "...", "speech_style": "..."
}],
"factions": [{ "id": "...", "name": "...", "motive": "..." }],
"truth": { "public": "...", "revealed": ["..."], "hidden": "..." },
"chapters": [{ "id": "...", "title": "...", "summary": "...", "loc": "...", "start_event": "..." }],
"clues": [{ "id": "...", "name": "...", "from_event": "..." }],
"endings": [{ "id": "...", "name": "...", "min": 0, "max": 100 }],
"events": [
{
"id": "...",
"title": "...",
"summary": "...",
"chapter_id": "...",
"preconditions": ["..."],
"dialogue": [{ "npc": "...", "line": "..." }],
"check": { "on_pass": "...", "on_fail": "...", "fail_cost": "..." },
"choices": [{
"id": "...", "text": "...", "next": "...",
"cost": "...", "effects": "...", "delta": -15, "payoff_at": "..."
}]
}
]
}| Category | Technology |
|---|---|
| Language | Python 3.12 |
| Vector store | Chroma (embedded PersistentClient, local folder) |
| Embedding | bge-m3 (FlagEmbedding, local, offline, no API key) |
| Multi-agent framework | CrewAI |
| LLM routing | OpenRouter (via CrewAI's LLM + litellm's openrouter/ prefix) |
| Validation | pydantic |
Supported environments: Python ≥ 3.12 (crewai requires ≥ 3.10; this repo standardizes on 3.12).
This project's code is licensed under the MIT License.

