Skip to content

About

PySide6 translation editor for Bethesda Starfield — edits .strings/.dlstrings/.ilstrings, BA2 archives and ESP/ESM plugins with Ollama or Claude AI backends, QC checks, translation memory, glossary, and NexusMods integration

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

413 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Bethesda Strings Editor

AI-assisted localization tool for Bethesda game files (Starfield). Translates .strings, .dlstrings, .ilstrings, BA2 archives, ESP/ESM plugin files, and Starfield interface TXT files between all 12 supported languages using a locally-running Ollama model, the Claude API, or your Claude Code subscription (via the bundled claude CLI — no per-token API cost), with a full quality-checking and review workflow.

NexusMods Header

Python PySide6 Ollama Claude License NexusMods


Ollama models

Model Purpose Hub
translategemma3-st Game string translation — Gemma 3 12B fine-tune (6.5 GB) local GGUF
translategemma3-st-2 Same fine-tune, reduced context (8 k) for GPU inference local GGUF
mamaylm MamayLM Gemma 3 12B IT v2.0 — INSAIT Ukrainian fine-tune local GGUF
gemma4-opus48-st Gemma 4 12B IT fine-tuned on Claude Opus reasoning data local GGUF
qcgemma4-st Translation quality checking — Gemma 4 E4B fine-tune (16 issue codes) 0xra/bethesda-qc

The custom translation fine-tunes are not published on any hub — there is no ollama pull for translategemma3-st. To translate, use your Claude Code subscription, the Claude API, or a local GGUF, as below.

Option A — Claude Code (your subscription, no API cost)

If you already pay for a Claude Pro / Max subscription, install Claude Code, run claude once to log in, then pick a Claude Code model (claude-code:haiku / claude-code:sonnet / claude-code:opus) in Settings → model. Translation, the chat assistant, and quality review shell out to the local claude CLI, so they run on your subscription — no Anthropic API key and no per-token billing. Each tier tracks the newest model your plan includes (e.g. sonnet currently resolves to Sonnet 5). No Ollama or GPU required.

Option B — Claude API (no local model)

Enter your Anthropic API key in Settings → Claude AI and pick Haiku 4.5 / Sonnet 4.6 / Opus 4.8. No Ollama or GPU required — the quickest way to get translating, but billed per token.

Option C — local Ollama model from a GGUF

Any Gemma-3 / instruct GGUF works with the app's prompts — the app supplies its own system prompt and overrides temperature, context size and output budget at runtime (the sampling knobs a Modelfile sets, such as min_p and repeat_last_n, stay in force; each Modelfile documents which). For Ukrainian, the publicly available MamayLM fine-tune is a good default:

# 1. Download a GGUF — e.g. MamayLM (Ukrainian) from HuggingFace:
#    https://huggingface.co/INSAIT-Institute/MamayLM-Gemma-3-12B-IT-v2.0
#    (or use any *.gguf you already have)

# 2. Edit the bundled Modelfile and point its FROM line at your file:
#    FROM /path/to/your-model.Q4_K_M.gguf

# 3. Build the model, then select it in Settings → Ollama:
ollama create translategemma3-st -f Modelfile

Modelfile, Modelfile.mamaylm, and Modelfile.gemma4-opus48 are templates — set the FROM path to your GGUF before running ollama create. The created model name is arbitrary; just select whatever you created in Settings → Ollama.

Optional — AI quality-check model (published on the hub)

ollama pull 0xra/bethesda-qc
ollama cp 0xra/bethesda-qc qcgemma4-st

Supported languages

All 9 official Starfield languages plus Russian, Ukrainian, and Korean:

Code Language
en English
de German
es Spanish
fr French
it Italian
ja Japanese
ko Korean
pl Polish
ptbr Portuguese (Brazil)
zhhans Chinese (Simplified)
ru Russian
uk Ukrainian

Each language pair has a dedicated system prompt with register rules, script conventions, and native examples.

Not listed here? A target language is an entry in a handful of tables, not a rewrite — FORKING.md walks through adding one, from the one-line minimum to the word lists and quality checks.


Features

Translation

  • Parallel AI translation via Ollama with configurable concurrency (default 10 workers)
  • Claude API backend — drop-in alternative to Ollama; select Haiku 4.5, Sonnet 4.6, or Opus 4.8 in Settings
  • Claude Code backend — run translation, chat, and review on your Claude Code subscription instead of the metered API, via the local claude CLI (claude-code:haiku / :sonnet / :opus); no API key, no per-token cost. The CLI's subscription (OAuth) auth is used and ANTHROPIC_API_KEY is stripped from the subprocess so it can never silently fall back to API billing
  • Rate-limit resilience — the Claude API client retries 429/5xx responses with exponential backoff (honouring Retry-After), so a large batch rides out transient rate limits instead of failing rows
  • AI-fix mode — instead of retranslating from source, sends the existing flawed translation + QC issue descriptions to the model for targeted correction; faster and more precise than a full retranslation
  • Language-pair prompts — dedicated system prompts for every source→target combination with register rules, script conventions, and native examples
  • Player-gender-aware translation — English "you" carries no grammatical gender, but most targets do, so the model would otherwise guess one per line. Declare the player character's gender once (Settings → Translation Preferences, or the Prompt Editor) and every backend applies it consistently; Translation → Find Player-Referring Strings selects the gender-sensitive rows. A pre-batch warning catches an unset gender before any AI call — and only when it matters: the target inflects and the batch actually contains player-referring lines
  • Translation Prompt Editor (Translation → Translation Prompt Editor) — customize the system prompt without editing code: set tuning dials (language style, formality, vocabulary, grammar, expression/localization, translation rigor), override the per-language style/register rule, and append project-wide instructions, with a live preview of the fully-assembled prompt. Applies to every backend (Ollama, Claude API, Claude Code CLI); the formatting-token protection rules stay fixed so placeholders can't be broken
  • Translation memory — known strings are looked up before calling the model, so they are never retranslated; a loaded TM is persisted as a JSON snapshot and auto-loaded next session, shown by a status-bar indicator, and inspectable via a searchable TM browser (Translation → Browse Translation Memory). Fuzzy matches whose numbers differ from the source (28LY ≠ 30LY) are rejected, and lookup stays fast on a six-figure TM via a sound pre-filter that only drops candidates the scorer would have rejected anyway
  • Official terminology miner (Translation → Mine Official Terminology) — Bethesda ships every official language for a plugin side by side on identical string IDs, so aligning English against an official target yields their canonical rendering of every weapon, faction, UI verb and quest term, with zero AI calls. Auto-detects the languages your install ships, mines a TM plus a filtered glossary (with optional reference languages — e.g. Polish as a Slavic cross-reference for Ukrainian), and imports into pending rows only, never clobbering work in progress. EN→DE on a real install: 190,367 TM entries + 17,815 glossary terms in 27 s
  • Apply to All Identical Originals (Ctrl+Alt+D) — propagate one row's translation to every row with the same source text in one shot; Delete clears a translation and reverts the row to pending
  • Translation cache — SHA-256-keyed JSON cache (up to 100,000 entries) persisted across sessions
  • Term protector — anything that must survive the model byte-for-byte is swapped for a placeholder token before the AI sees the text and restored afterward: 20 structural patterns (game tags, format specifiers, FormIDs, chemical formulas, and deliberately-obfuscated in-game codes such as encrypted-note passwords), plus a built-in list of 86 proper nouns. Starfield's faction, company and UI terms are meant to be translated, so no term list is bundled; drop a protected_terms_starfield_hq.txt next to the application and it is read at startup, for names that must stay verbatim in every locale
  • Pipeline post-processing — per-string passes after every translation: game tag restoration, case matching, line-prefix preservation, newline structure repair, mixed-script repair, guillemet close-quote enforcement
  • Glossary system — CSV/TBX/JSON glossary with in-app editor, term suggestions dock, and automatic injection into AI prompts
  • Character Persona Profiling — assign a voice profile to any string or quest (Freestar Ranger, SysDef Officer, Crimson Fleet Pirate, House Va'ruun Zealot, UC Civilian, Robot/Automaton, Narrator, or custom); each profile overrides the AI system prompt and temperature
  • Lore RAG — local SQLite FTS5 lore database (built-in UESP downloader); relevant faction, location, and character articles are retrieved per string and injected into the AI prompt
  • Pre-translation estimator — scores each string 0–100 to predict translation difficulty before the AI runs
  • Skip string types — exclude Book, Note, or other categories from AI batch translation
  • Protect named entities — opt-in setting to extend term protection to faction/ship/character names inferred from the loaded file
  • Claude pre-flight estimator — shows token count and estimated USD cost before starting a batch on the Claude API; on the Claude Code (subscription) backend it shows token counts only (no misleading cost) and reports the actual tokens used after the batch completes

File support

  • Binary string files: .strings (null-terminated), .dlstrings / .ilstrings (length-prefixed)
  • BA2 archives: read and write Starfield v2 BA2 files (GNRL type, zlib-compressed); picker dialog for multi-entry archives
  • ESP/ESM plugins: non-localized plugins where text is stored directly in field buffers; covers GameplayOption settings-menu titles (GPOF/GPOG), skips asset paths that masquerade as text (e.g. DOOR/CNAM animation paths), and adds an anti-hallucination context note to quest titles (QUST/FULL)
  • ESP/ESM Mod Update Migration (Translation → Mod Update Migration) — diff an old and new version of a plugin keyed on FormID and carry existing translations forward to the updated release (xTranslator-style); risk-coloured 7-column diff with changed-only filter and CSV/HTML export
  • VMAD script-property analysis (Translation → Script Property Analysis) — parses the Papyrus script-property strings attached to records, classifies each value as translatable / review / locked (resource paths, event names, identifiers), and byte-splices only the edited values so unmodelled script structures survive untouched; works on both localized and non-localized plugins
  • Starfield interface TXT: translate_en.txt / translate_ru.txt key=value interface string files
  • xTranslator SST XML: import/export in xTranslator format (match by sID, fallback to source text)
  • Translation folder validator (Translation → Validate Translation Folder) — scans a finished translation against the game's Data folder (loose files and .ba2 archives) and flags every file and ID that will show an in-game <Error: Unknown lstring ID …>, before you launch the game; also surfaces cross-file-type ID contamination
  • Companion strings viewer (Translation → Companion Strings) — read-only view of a loaded .strings/.dlstrings/.ilstrings triplet with file-type and text/ID filters. The three files have independent ID spaces, so companions load as reference and are never merged into the file you save
  • Drag-and-drop file loading with format validation
  • NexusMods Translation Browser — search NexusMods for existing translation mods, browse their files, and import .strings/.dlstrings/.ilstrings directly as a Translation Memory or merge into the current file; zip, 7z, and rar archives are automatically extracted; free-account downloads via browser cookies (curl-cffi); account sign-in uses the official browser-based Single Sign-On flow (no API key to paste — the app is a registered NexusMods application)

Quality assurance

  • Quality checker — 20+ checks: missing/extra game tags, empty or untranslated strings, source-language leakage, English leak, suspicious length ratios, newline mismatches, truncated AI output, AI artifact prefixes, encoding failures, unclosed guillemets, unmatched brackets, script coverage (CJK), and more
  • RU→UK false-positive reduction — UNTRANSLATED check uses Russian-exclusive character detection (ы/э/ё/ъ) and minimum word-count threshold to avoid flagging legitimately identical short words; length-ratio skip applies only to short Cyrillic sources (abbreviation expansions)
  • Hunspell spell-check — per-language SPELL_ERROR warnings; uses system dictionaries on Linux and app-bundled dictionaries on Windows/macOS (populate via scripts/fetch_dictionaries.py)
  • AI quality model (qcgemma4-st) — fine-tuned Gemma 4 E4B that detects 16 issue codes with chain-of-thought reasoning and structured VERDICT: GOOD / ISSUES_FOUND output with AUTOFIX/RETRANSLATE recommendations
  • Font & Glyph Checker — parses Scaleform SWF font atlases and TTF/OTF cmap tables; flags translation characters that will render as squares in-game and suggests auto-fixable substitutes
  • UI Width-Fit Simulator (Translation → UI Width-Fit Simulator) — the glyph checker proves a character exists; this proves the label it spells fits its box. Cyrillic, German and Polish run 15–30 % longer than English and Scaleform widgets clip rather than shrink, so a perfectly-spelled translation still ships as a truncated button. Width is summed from the real per-glyph advances of the actual game faces (SWF FontAdvanceTable / TTF hmtx), with markup resolved and value placeholders substituted first. Results are colour-coded worst-overflow-first with CSV export, and every row also carries vs Source — the translated width over the English width, which needs no budget at all, since the English fit by construction
  • Real widget bounds, read from the game's own SWFs — budgets are not guessed. Point the simulator at your Data folder and it reads every Interface/*.swf and *Interface*.ba2 (~250 SWFs, ~2560 clipping fields), taking each field's authored width, margins, clip behaviour and font class straight from its DefineEditText record. Font size is marked declared when the SWF states it and derived when inferred from box height — never flattened together, because width scales linearly with size, so a 20 % size error is a 20 % wrong verdict
  • Large-font (accessibility) menus checked as the worst case — Starfield ships a _lrg build of most menus where the box usually stays the same size while the font grows (up to ×2.64), so a label can pass the standard menu and clip only for players using the large font. The tightest of the two builds is chosen per widget, and the dialog always names which build it measured — including when the twin could not be identified, so "no warning" never quietly means "not checked"
  • Korean particle (조사) checker — allomorphs are computed from Hangul 받침 arithmetic, never guessed. Flags single-form particles after a value placeholder (always a latent bug: the substituted noun's final consonant is unknowable at translation time and the engine resolves nothing), and particle/stem mismatches within the provably sound subset. Measured at 0 false positives over 5,353 real Korean strings; both codes are auto-fixable, and the model is taught the same rule
  • Auto-Fix All — one-click batch application of all mechanically correctable issues (whitespace, capitalization, character substitution, missing newlines, truncated translations, unclosed guillemets)
  • Per-code hide filter — suppress specific QC issue codes from the results table for the current session
  • Retranslation queue — strings flagged by QC are queued and retranslated with a per-string hint describing what went wrong
  • Error-code filter — filter QC results by code (MISSING_TAGS, NEWLINE_COUNT_MISMATCH, etc.)
  • Consistency checker (Ctrl+Alt+K) — finds the same source string translated differently across the file, with canonical-form picker and batch replace
  • Ukrainian gender agreement checker (Ctrl+Alt+G) — detects adjective/noun gender mismatches using a noun-gender dictionary, with inline fix suggestions
  • ти/ви register consistency — informal ти / formal ви is kept consistent per speaker inline by the translation system prompt (shared by the Ollama and Claude backends), so mixed address never enters the output in the first place
  • Plugin validator — scans ESP/ESM for NPC dialogue camera bugs: missing Localized flag, stray DIAL/SCEN/INFO records, ONAM overrides, missing master dependencies

Review tools

  • Visual Context Preview (Ctrl+Shift+P) — dockable panel that renders the current string inside a faithful recreation of the Bethesda UI using actual game fonts extracted from fonts_uk.swf / fonts_en.swf; pixel-exact borders, noise tile, and dark gradient from dialoguemenu.swf; auto-detects context type (Dialogue, Quest, Book, Note, Terminal, UI); Source/Translation/Both view modes. Carries a live fit indicator for length-critical strings (Tab label: 262/220px (119%) CLIPS) that measures real font advances against a fixed widget budget — deliberately separate from the canvas overflow badge beside it, which only reports whether the text outgrew the preview, and so cannot answer "will this clip in-game?"
  • Dialogue Tree Visualizer — interactive quest → topic → response node graph (Translation → Dialogue Tree) with Starfield dark-space visual theme; click any node to jump to that string in the table
  • Audio / TTS Preview (Ctrl+Shift+A) — dockable panel with eSpeak-NG and Piper backends; synthesizes a TTS read-out of the translation so timing can be compared against the original game audio; colour-coded timing bar (green ≤ 110 %, orange ≤ 130 %, red > 130 %)
  • Native Starfield voice playback — decodes the original Wwise .wem voice clip for a dialogue line straight from the game's *Voices*.ba2 archives (FormID → voice clip, via vgmstream-cli) and plays it back in the audio panel alongside the TTS read-out; in ESP/ESM mode the row's FormID auto-fills
  • Speaker (NPC) panel — shows who voices the selected dialogue line (name, gender, faction, category, raw voice type, plus "also voiced by" for shared lines) by parsing the Wwise voice-type folder name; tabified with the Audio panel
  • Version comparison — diff two game versions, migrate unchanged translations, export CSV/HTML reports; batch folder comparison
  • Diff viewer — side-by-side word-level or character-level diff; editable right pane with live diff update; HTML export
  • Advanced search — regex and fuzzy search across source and translation columns; batch Find & Replace

UI / workflow

  • Zen / Focus Mode (F11) — full-screen distraction-free editor with large source and translation panels, pending-string counter, per-string status badge
  • Multi-monitor / detached panes — Translation Editor dock (Ctrl+Shift+E) floats to any monitor; Pop Out String List (Ctrl+Shift+L) opens a second table window sharing the same selection model
  • Claude AI Assistant dock (Ctrl+Shift+C) — chat about the current string and apply Claude's suggested translation with one click; optionally connect remote MCP servers (Settings → Claude MCP Servers) so Claude can call external tools while you chat, via the Messages API MCP connector
  • Command palette (Ctrl+K) and vim-style navigation (j/k, G)
  • Translation sessions — named work sessions with persistent search/filter state (Ctrl+Shift+N new, Ctrl+Shift+S save)
  • Macro recorder (Ctrl+M) — record and replay sequences of edit operations as named macros
  • Keyboard shortcuts editor — rebind any action
  • F7 → jump to next untranslated; Ctrl+Enter → approve; Ctrl+R → reject
  • Encoding detection — auto-detects UTF-8, CP1251, CP1252, CP1250, GBK/GB2312, BOM variants; override per-file
  • Themes — 16 built-in themes: Slate, Midnight, Nord, Dracula, Catppuccin, Light, Solarized Dark, Solarized Light, Gruvbox, Tokyo Night, Monokai, One Dark, Sepia, Starfield, Starfield Terminal, High Contrast; plus custom QSS file support
  • GPU monitor — status bar widget showing GPU utilisation, VRAM usage, and temperature (AMD via Linux sysfs; NVIDIA via nvidia-smi on Linux/Windows/macOS; auto-hides if no GPU found)
  • Ollama model auto-detection — the model dropdown in Settings loads installed models automatically and keeps refreshing while the window is open, so a model pulled with ollama pull shows up without clicking Refresh
  • Ollama force-stop — frees a wedged GPU by restarting the Ollama service; on Linux a privileged restart uses the app's own Qt-themed password dialog (sudo -S, with askpass/pkexec fallback), on Windows it stops the service via taskkill with no console flash
  • UI translations — interface available in Ukrainian ✓, German, Spanish, French, Polish, Czech, Korean (community-maintained, machine-assisted; native-speaker review welcome — see TRANSLATING.md)
  • Cross-platform desktop notifications on batch completion — notify-send/D-Bus on Linux, native system-tray balloons on Windows and macOS
  • Native OS integration — native Explorer/Finder file dialogs on Windows/macOS; config stored in the OS-native location (%APPDATA% on Windows, ~/Library/Application Support on macOS, $XDG_CONFIG_HOME/~/.config on Linux) with owner-only file permissions
  • "What's New" panel — recent GitHub release notes are fetched and rendered on the welcome screen after the welcome card
  • Automatic update check — checks the GitHub releases API on startup and offers to download a newer build (toggle in Settings)
  • Crash recovery — periodic auto-save; recovery dialog offered at startup if the previous session ended unexpectedly
  • Security audit log — append-only JSON-lines file recording file operations and translation batches; API keys stored in system keyring with AES-256-GCM file fallback

Requirements

  • Python 3.10+
  • One AI backend: Ollama running locally, or a Claude API key, or the Claude Code CLI logged in to your Pro/Max subscription (no API cost)
  • Audio playback is auto-detected per platform: Linux uses paplay, ffplay, or aplay; macOS uses afplay (or ffplay); Windows uses ffplay or the built-in PowerShell WAV player
  • Native Starfield voice playback requires vgmstream-cli on PATH (decodes Wwise .wem clips)
pip install -r requirements.txt

Core dependencies: PySide6>=6.11, requests>=2.33, cryptography>=48.0, anthropic>=0.104

Optional — each has a runtime fallback, so install only what you need:

  • keyring>=25.7 — API keys in the system keyring (otherwise an AES-256-GCM file store is used)
  • curl-cffi>=0.15 — free-account NexusMods downloads via browser cookies
  • py7zr>=1.1 / rarfile>=4.3 — .7z and .rar extraction (otherwise the 7z / unrar CLIs)
  • hunspell>=0.5.5 or spylls>=0.1.7 — spell-check engine (hunspell CLI used as fallback); run scripts/fetch_dictionaries.py to bundle Hunspell dictionaries for Windows/macOS

Developed and tested on Python 3.10 with the versions above; the test suite runs green on them.


Running

python main.py

Logs are written to both stdout and translator.log in the project root.

File associations (Linux and Windows)

Register the app as the handler for .strings / .dlstrings / .ilstrings, .esp / .esm / .esl and .ba2 — file icons, the Open With entry, and double-click:

bethesda-strings-editor --register-file-types      # frozen build
python main.py --register-file-types               # from source
python main.py --unregister-file-types             # undo

Everything is per-user and needs no root or admin: $XDG_DATA_HOME on Linux (MIME XML, desktop entry, icons in both the apps and mimetypes contexts, Thunar restarted so it drops its icon cache), HKCU\Software\Classes on Windows. On Windows the app is added to Open With but does not claim the default handler unless you pass --force, since Windows guards a default you chose yourself and quietly taking .esp from another modding tool would be worse than not being the default. Linux users installing from a checkout can use scripts/install_file_associations.sh instead — it is a wrapper around the same code.


Project structure

bethesda_strings/              Pure Python parsing library (no Qt dependency)
  core.py                      Binary parser/writer for .strings/.dlstrings/.ilstrings
  ba2_handler.py               BA2 archive reader/writer (Starfield v2, FO4 v1)
  esp_handler.py               ESP/ESM plugin parser (non-localized plugins)
  esp_diff.py                  ESP/ESM mod-update migration (FormID-keyed diff)
  vmad_handler.py              VMAD Papyrus script-property parser/classifier/byte-splice editor
  wwise_voice.py               Wwise voice index — FormID → original .wem clip from Voices BA2
  txt_handler.py               Starfield interface TXT parser (translate_en/ru.txt)
  xml_handler.py               xTranslator SST XML import/export
  encoding.py                  Encoding detection and conversion
  version_diff.py              Game-version diff and translation migration
  triplet.py                   Read-only .strings/.dlstrings/.ilstrings companion holder (independent ID spaces)
  strings_validator.py         Classifies a translation folder against the game's source IDs
  official_tm_miner.py         Mines TM + glossary from the base game's own official localizations
  character_profiles.py        Character persona profile definitions
  font_checker.py              SWF/TTF glyph coverage + advance-width metrics (library layer)
  swf.py                       SWF container primitives (decompress, tag walk, bit-packed RECT)
  swf_widgets.py               Real UI widget bounds from the game's DefineEditText records
  width_fit.py                 UI width-fit simulator — measured width vs. widget budget
  lore_db.py                   SQLite FTS5 lore database and UESP downloader
  dialogue_tree.py             Quest → Topic → Response tree parser

gui/                           PySide6 application layer
  main_window.py               Top-level window, file I/O, translation orchestration
  ollama_worker.py             QThread worker — parallel calls, per-language prompts, AI-fix mode
  ollama_control.py            Ollama service force-stop/restart helper (sudo/taskkill)
  sudo_dialog.py               Qt-themed sudo password prompt (sudo -S over stdin)
  claude_translation_worker.py Claude API / Claude Code drop-in replacement for OllamaWorker
  claude_code_client.py        Claude Code CLI backend (subscription, no API cost) — ClaudeClient drop-in
  claude_chat_panel.py         Dockable AI assistant chat panel
  gpu_monitor.py               Status bar GPU utilisation widget (AMD sysfs + NVIDIA nvidia-smi)
  visual_context_preview.py    Game-accurate string rendering using extracted SWF fonts/assets
  dialogue_tree_dialog.py      Interactive quest → topic → response node graph
  audio_preview_panel.py       TTS preview dock + native Wwise voice playback (eSpeak-NG / Piper)
  tts_engine.py                TTS synthesis engine
  speaker_map.py               Wwise voice-type → speaker (name/gender/faction) parser
  speaker_panel.py             Speaker (NPC) dock — who voices the selected line
  focus_overlay.py             Zen / full-screen focus mode
  lore_rag_manager.py          SQLite FTS5 lore database + UESP downloader
  lore_rag_dialog.py           Lore database management dialog
  quality_checker.py           Post-translation QA checks (20+ codes) and auto-fix
  quality_dialog.py            QA results dialog — filtering, per-code hide, auto-fix, retranslation
  ai_qc_worker.py              Worker thread for qcgemma4-st quality model
  spell_checker.py             Hunspell spell-check wrapper (3 backends: lib / spylls / CLI)
  ko_particle_checker.py       Korean 조사 agreement, computed from Hangul 받침 arithmetic
  player_gender.py             Detects player-referring (gender-sensitive) source strings
  font_checker_dialog.py       SWF/TTF glyph coverage checker dialog
  width_fit_dialog.py          UI width-fit simulator UI (real widgets, large-font worst case)
  string_table.py              QAbstractTableModel for strings, ESP, and TXT modes
  term_protector.py            Placeholder-based term protection (8000+ terms)
  translation_cache.py         SHA-256-keyed persistent translation cache
  translation_memory.py        Pre-loaded map of string ID → known-good translation (+ JSON snapshot)
  tm_browser_dialog.py         Searchable read-only Translation Memory browser
  official_tm_dialog.py        Official-terminology miner UI (language auto-detect, glossary preview)
  fuzzy_match.py               Fuzzy scoring + FuzzyIndex candidate pre-filter for TM lookup
  validate_translation_dialog.py  Translation-folder validator UI (in-game lstring-ID errors)
  companion_strings_dialog.py  Read-only viewer for a loaded .strings triplet
  glossary.py                  Glossary data model, CSV/TBX/JSON I/O
  consistency_checker.py       Finds inconsistent translations of identical source strings
  gender_checker.py            Ukrainian adjective/noun gender agreement checker
  prompt_editor_dialog.py      Translation system-prompt editor (per-language style rule + addendum, live preview)
  vmad_dialog.py               VMAD script-property analysis UI (risk-coloured, byte-splice apply)
  esp_migrate_dialog.py        ESP/ESM mod-update migration UI (FormID-keyed diff)
  version_compare_dialog.py    Game-version diff UI, migration, CSV/HTML export
  diff_viewer.py               Side-by-side word/character-level diff viewer
  pre_translation_estimator.py Difficulty scorer (0–100) with weight learning
  profile_editor_dialog.py     Character persona profile editor
  profile_assign_dialog.py     Assign persona profiles to strings / quests
  keyboard_manager.py          Rebindable shortcuts, vim navigation, command palette
  session_manager.py           Named work sessions with persistent search/filter state
  macro_recorder.py            Record/replay sequences of edit operations as macros
  theme_manager.py             16 built-in QSS themes + custom theme loader
  nexusmods_uploader.py        NexusMods v3 multipart upload client (used only by the release CI)
  nexusmods_browser_dialog.py  NexusMods translation browser (card grid, async thumbnails)
  nexusmods_client.py          NexusMods REST v1 + GraphQL v2 API client
  app_settings.py              AppSettings dataclass, JSON + QSettings persistence
  secret_store.py              API key storage (keyring + AES-256-GCM fallback)
  audit_log.py                 Append-only security audit log (JSON-lines)
  updater.py                   GitHub release update check + welcome "What's New" changelog
  desktop_notify.py            Cross-platform batch-complete notifications (notify-send / tray)
  crash_recovery.py            Periodic auto-save and recovery dialog

data/                          (see data/README.md for provenance and exact figures)
  fonts/                       Game fonts exported from Starfield's Scaleform SWFs with JPEXS
                               (also the advance-width source for the width-fit simulator).
                               All five are 1024 units/em and declare weight 400 — the weight
                               is in the outlines, so faces are picked by family, not metadata
    RF_35_M.ttf                Cyrillic body face (UK locale, $MAIN_Font); 573 glyphs
    RF_55_M.ttf                Cyrillic bold role — ~10 % wider per glyph than RF_35_M
    RF_55_SB.ttf               Cyrillic semibold — ~16 % wider, ~19 % on the ALL-CAPS
                               text that button labels are made of
    NB_Architekt_Light.ttf     Latin display face (EN locale); no Cyrillic coverage at all
    NB_Architekt.ttf           Latin bold role — only ~3 % wider than the Light cut
  dialogue_bg_tile.png         50×50 noise tile from dialoguemenu.swf (tiled by the preview)
  dialogue_panel_ref.png       597×147 export of the panel sprite the preview's measurements
                               were verified against; reference only, no code loads it
  *_words.txt                  Word lists for source-language leak detection (10 languages) and,
                               for Korean, the 조사 particle checker; loaded per language pair
                               rather than all at once (~329 MB → ~45 MB for an EN→KO session)

scripts/
  extract_sharegpt_dataset.py  Export EN→target string pairs as ShareGPT JSONL
  create_qc_dataset.py         Generate QC training dataset (14,928 examples, 16 issue codes)
  compile_translations.sh      Recompile .ts → .qm UI translation files
  fetch_dictionaries.py        Download Hunspell .aff/.dic dictionaries into dicts/ (Windows/macOS spell-check bundle)
  download_lang_dicts.py       Download word-frequency lists into data/ (de/es/fr/it/ko/pl/ptbr;
                               reproduces the committed files byte for byte)
  extract_starfield_glossary.py Build starfield_glossary.json from string files (local artifact, untracked)
  nexusmods_upload.py          Publish a release to NexusMods (driven by the release workflow)
  install_file_associations.sh Register the Linux desktop entry, MIME types and icons

packaging/
  bethesda-strings-editor.desktop      Linux desktop entry (file associations)
  bethesda-strings-editor-mime.xml     MIME-type definitions for .strings/.esp/.ba2

UI translation

UI translations live in gui/translations/<locale>.ts. After editing any .ts file:

./scripts/compile_translations.sh

Supported locales: uk_UA, de_DE, fr_FR, es_ES, pl_PL, cs_CZ, ko_KR — 1,735 strings each, all filled in.

English and Ukrainian are maintained by @0xra0; every other locale is community-maintained — machine-assisted and not proofread by a native speaker, because the author does not speak them. Corrections are very welcome: see TRANSLATING.md.

Note that this is the app's own UI, which is a different thing from the language it translates game strings into — see FORKING.md for that one.


Forking

FORKING.md is the guide for making this editor your own:

  • Adding a translation target language — the one-line minimum, then display name, style rule, encoding pair, gender, word list + checker, spell-check and quality-check wiring, with a checklist and a worked Czech example
  • Custom prompts — the Prompt Editor's three layers with no code, and how to_system_prompt() assembles them with code (including which parts of the built-in prompt are Ukrainian-specific and should be swapped in a fork)
  • Adding a translation backend — the worker signal interface and where routing happens
  • Other extension points — themes, plugin field types, string categories, lore injection, new settings
  • Keeping a fork mergeable with upstream

Everything there is an entry in an existing table rather than a modified line, which is what lets a fork rebase cleanly.


Tests

python -m pytest tests/

1,007 tests across 51 files, all pure Python — no Qt event loop, no network, no game install. Binary formats are exercised with hand-built record and tag bytes and the AI backends with fake clients; the one end-to-end voice-decode test self-skips when Starfield and vgmstream-cli aren't present.


License

MIT — see LICENSE.

About

PySide6 translation editor for Bethesda Starfield — edits .strings/.dlstrings/.ilstrings, BA2 archives and ESP/ESM plugins with Ollama or Claude AI backends, QC checks, translation memory, glossary, and NexusMods integration

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages