Skip to content

Latest commit

 

History

History
193 lines (155 loc) · 8.76 KB

File metadata and controls

193 lines (155 loc) · 8.76 KB

BiXiaScribe

Retrieve wuxia source text with RAG, hand it to a multi-agent pipeline that writes a structured wuxia RPG script.

License: MIT Python Stars

繁體中文 | English


What is this?

BiXiaScribe is a wuxia (武俠) RPG script generator. Give it one line of a story requirement (e.g. "a Shaolin disciple leaves the temple to investigate a massacre"), and it retrieves relevant passages from your own corpus of wuxia novels, then hands them to three specialized LLM agents (writer → dialogue → proofreader) that produce a structured script JSON — NPCs, events, branching choices, triggers — usable as source material for downstream game production (e.g. RPG Maker).

UI preview

A browser UI for reading and comparing scripts already generated by scripts/eval_generation.py, instead of opening out/eval/*.json by hand. Four modes — single-script / side-by-side / overview table / generate. The generate mode triggers a real run from the browser on a background thread, with a live elapsed-time clock, a task-boundary progress bar, and a cancel button that actually interrupts the run.

Single-script view - events tab Events tab: triggers, dialogue lines, and branching choices rendered as readable prose instead of raw JSON.

NPC tab: the character table (id / name / identity / personality / speech style). Run tab: the joined RunReport for this generation — per-agent model ids, elapsed time, retrieval_calls, repair_attempts, total_tokens, coerced_from.

Surfacing retrieval_calls per script in the UI is what makes the zero-retrieval-call finding from Key results below checkable at a glance, instead of something you have to infer from log scrollback.

Features

  • RAG indexing: Chinese-aware chunking → embedding → Chroma, resumable.
  • Hybrid retrieval (vector + hand-rolled BM25, see Key results below).
  • Layered/stateful script generation pipeline (extract → beats → scene-by-scene → proofread, real-time causal-graph validation, checkpointed resume, batch confirmation), structured JSON + cross-reference validation.
  • Per-agent model-split A/B and cost accounting (scripts/eval_generation.py); unit tests never hit a real API.
  • Retrieval can be switched off entirely (RETRIEVAL_ENABLED=false / --no-retrieval / a UI checkbox) to save the single largest token cost, for A/B'ing whether a model with a strong native wuxia voice actually needs corpus grounding.
  • Streamlit UI, four modes including browser-triggered generation — see UI preview above. Can delete/export/import scripts, restricts model selection to a curated catalog, sets a global reasoning-effort knob, and resumes an interrupted layered run (hard-gated on checkpoint schema version).
  • 📋 Planned: editing script content and saving it back from the UI.

Why this architecture for wuxia scripts

Compared to just prompting ChatGPT directly for a script, BiXiaScribe differs in:

  • RAG retrieval over real source text, not just the model's imagination of "wuxia voice" — it indexes your own corpus and feeds retrieved passages into the dialogue agent's prompt, so wording and move names stay closer to the source material.
  • A Chinese-aware chunker — measures length in characters and prefers splitting at paragraph/punctuation boundaries, instead of reusing token-splitting logic built for English NLP.
  • Structured output with automated cross-reference validation — dialogue.npc/choices[].next cross-references are re-checked in Python, not just trusted because the proofreader agent says it's fine.
  • Local-first, runnable end-to-end at zero cost — the default embedding backend is the local bge-m3 model (offline, free, no API key); the LLM also has a fake mode so tests never make a real API call.

Key results

  • Retrieval: under a strict comparison (--top-k 1), term-match rate is 91.7% for hybrid vs. 75% for vector-only — character-bigram BM25 catches proper-noun matches vector search tends to miss.
  • Generation no-RAG A/B (2026-08-19): the retrieval-on variant's usd_per_event ($0.0043) is lower than the retrieval-off variant's ($0.0081) — despite costing 40% more per run — while also having a higher NPC-speaking rate and longer dialogue lines, and cutting dead-end self-loop branches from 20% to 0%.
  • A still-valid non-obvious finding: retrieval_calls shows that "the model supports function calling" per OpenRouter's metadata is not the same guarantee as "it reliably chooses to call the tool in a CrewAI ReAct loop" — this needs checking per model.

Full tables, methodology, and past A/B rounds: docs/BENCHMARKS.md (in Chinese).

Quickstart

python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # only needed to run real script generation (OpenRouter)
# 1. Build an index (sample corpus, done in seconds, no API key)
python scripts/build_index.py --corpus tests/sample_corpus.txt

# 2. Query the index (default: hybrid = vector + BM25)
python scripts/test_retrieval.py --query "獨孤九劍的劍法精要" --top-k 3

# 3. Check the backend/API key/index are wired up before spending a token
python scripts/generate_script.py --requirement "test" --preflight-only

# 4. Generate a script (needs LLM_BACKEND=openrouter + OPENROUTER_API_KEY) --
#    checkpointed under .bixia_state/<run_id>/, resumable (see CLAUDE.md)
python scripts/generate_script.py --requirement "少林弟子下山查一樁滅門案" --out script.json
python scripts/generate_script.py --requirement "..." --run-id <run_id>   # resume a specific run

# 5. Browse/compare already-generated scripts (no API key, no tokens spent),
#    or use the 生成 (generate) mode to trigger a real run from the browser
pip install -r requirements-ui.txt
.venv/bin/streamlit run ui/app.py

(screenshots above, see UI preview)

Drop your own .txt files into data/corpus/ (not assumed UTF-8). Full commands for switching corpora/embedding backends and comparing retrieval/model-split quality are in docs/DESIGN_NOTES.md (in Chinese).

Output format

script.json looks roughly like this (full field definitions in src/bixiascribe/schema.py):

{
  "meta": { "title": "...", "theme": "...", "goal": "...", "tone": "..." },
  "stat": { "id": "mood", "name": "mood value", "init": 50 },
  "player": { "name": "...", "origin": "...", "flaw": "...", "token": "..." },
  "items": [{ "id": "...", "name": "...", "from_event": "..." }],
  "npcs": [{
    "id": "...", "name": "...", "faction_id": "...", "role": "...",
    "personality": "...", "speech_style": "..."
  }],
  "factions": [{ "id": "...", "name": "...", "motive": "..." }],
  "truth": { "public": "...", "revealed": ["..."], "hidden": "..." },
  "chapters": [{ "id": "...", "title": "...", "summary": "...", "loc": "...", "start_event": "..." }],
  "clues": [{ "id": "...", "name": "...", "from_event": "..." }],
  "endings": [{ "id": "...", "name": "...", "min": 0, "max": 100 }],
  "events": [
    {
      "id": "...",
      "title": "...",
      "summary": "...",
      "chapter_id": "...",
      "preconditions": ["..."],
      "dialogue": [{ "npc": "...", "line": "..." }],
      "check": { "on_pass": "...", "on_fail": "...", "fail_cost": "..." },
      "choices": [{
        "id": "...", "text": "...", "next": "...",
        "cost": "...", "effects": "...", "delta": -15, "payoff_at": "..."
      }]
    }
  ]
}

Tech Stack

Category Technology
Language Python 3.12
Vector store Chroma (embedded PersistentClient, local folder)
Embedding bge-m3 (FlagEmbedding, local, offline, no API key)
Multi-agent framework CrewAI
LLM routing OpenRouter (via CrewAI's LLM + litellm's openrouter/ prefix)
Validation pydantic

Supported environments: Python ≥ 3.12 (crewai requires ≥ 3.10; this repo standardizes on 3.12).

License

This project's code is licensed under the MIT License.