⚠ AI-Generated / Work in Progress
The code, lesson packages, exercises, Anki and Flashcards Deluxe decks, and analysis outputs in this repository were generated by Claude (Anthropic) under human direction. This is an experimental project and is actively being developed. You may encounter errors, incomplete content, inaccurate linguistic analysis, or other rough edges.
Feedback and bug reports are welcome — please open a GitHub issue.
A Python project for morphological analysis, statistical reporting, and language instruction across the full biblical corpus — Hebrew and Aramaic Old Testament, Greek New Testament, Greek Septuagint (LXX), Syriac Peshitta NT, Aramaic Targumim, and major English translations.
The project runs on three tracks simultaneously:
- Research tool — 50+ analysis functions that answer cross-corpus morphological and semantic questions
- Teaching platform — lesson packages for Biblical Hebrew (BBH), Greek (BBG), and Aramaic (BBA), with interactive exercises, Anki decks, and live course sessions
- MCP server — a semantic search interface over a curated biblical-studies research library, usable directly inside Claude
Everything publishes to bereanbiblebots.com — a static MkDocs site with full lesson content, interactive exercises, course session pages, analysis reports, and Jupyter notebooks runnable on Google Colab.
- How many Niphal perfect verbs are in each book of the OT?
- What is the verb stem distribution across the Torah?
- How does Paul use the aorist passive compared to the rest of the NT?
- Where does the word "grace" appear in Paul's letters (KJV)?
- How consistently does the LXX render רוּחַ (spirit/wind) across books?
- Does Hebrews quote the OT following the LXX or the Hebrew MT?
- What Greek word does the LXX use to translate חֶסֶד (lovingkindness), and how does that word travel into the NT?
- Which words cluster significantly near שָׁלוֹם (peace) in the OT?
- How does Paul's use of Χριστός (Christ) compare to the Gospels?
- How do verb conjugation patterns differ between narrative prose and wisdom poetry?
- How does Targum Jonathan render Isaiah 53?
# 1. Clone with submodules
git clone --recurse-submodules https://github.com/dnovick/berean-bible-bots.git
cd berean-bible-bots
# 2. One-command setup — creates a virtual environment, installs all dependencies,
# and registers the Jupyter kernel
bash setup.sh # Mac / Linux
# setup.bat # Windows
# 3. Build the processed database (one-time, ~5 seconds)
source .venv/bin/activate # Mac / Linux
# .venv\Scripts\activate # Windows
python scripts/build_db.py
# 4. Download Targum data from Sefaria (one-time, ~14,000 verses)
python scripts/fetch_targum_data.py
# 5. Open any notebook in notebooks/ — or run on Google Colab (no install needed)Google Colab (no local install): every notebook at bereanbiblebots.com has an Open in Colab badge. The first cell clones the repo and downloads the processed data (~295 MB, one-time per session).
See docs/getting-started.md for full developer setup and API usage.
The project includes a research-library MCP server that exposes a curated collection of biblical-studies reference works — grammars, commentaries, lexicons, and journal articles — as a semantic search tool usable directly inside Claude.
Setup:
# 1. Index the research library (one-time, requires OpenAI API key)
python scripts/index_research_library.py
# 2. Add the server to your Claude Code config (.mcp.json)
# See .mcp.json in the repo root for the configuration templateTools exposed:
| Tool | Description |
|---|---|
search_literature(query, top_k=5) |
Semantic search over indexed grammars, commentaries, and journal articles — returns ranked excerpts with citation metadata |
get_indexed_sources() |
Lists all works in the index with title, author, year, tags, and chunk count |
The server uses ChromaDB with OpenAI embeddings. Once indexed, it can be queried from within any Claude conversation to pull primary-source support for exegetical or grammatical questions.
Structured lessons for three biblical language courses, all published at bereanbiblebots.com:
| Course | Chapters | Textbook |
|---|---|---|
| BBH — Biblical Hebrew | Ch1–35 (alphabet → Hithpael weak) | Basics of Biblical Hebrew, Pratico & Van Pelt |
| BBG — Biblical Greek | Ch1–36 (alphabet → μι-verbs) | Basics of Biblical Greek, Mounce (4th ed.) |
| BBA — Biblical Aramaic | Ch1–22 (alphabet → Hophal stem) | Basics of Biblical Aramaic, Van Pelt |
Each chapter includes:
- Lesson page — full notes with paradigm tables, key terms, and examples
- Interactive exercises — fillable HTML with per-item answer reveal; AcroForm PDFs for printing
- Flashcard decks — morphology and vocabulary decks in Anki (
.txt) and Flashcards Deluxe (-fd.txt) formats
Active course sessions (BBH 2024.1 and BBH 2026.1) are tracked in data/courses/ and
published as session pages with instructor notes, session content, and linked resources.
See docs/lesson-packages.md for exercise formats and PDF generation details.
- Filtered morphological access to Hebrew OT, Greek NT, LXX, KJV, and Vulgate
- Word study and semantic profile — lexicon, frequency, morphology, collocations
- Cross-testament trajectory — OT → LXX → NT word journey with continuity scoring
- NT quotations and word alignment — LXX vs. MT source tracking for every NT citation
- Verbal syntax — wayyiqtol chains, conditionals, relatives, aspect comparison, discourse particles
- Poetry analysis — cola splitting, parallelism, chiasm, acrostic detection, meter
- Noun and number morphology — state/gender/number profiles; gender-polarity rule for cardinals
- Predicate-argument structure — PropBank A0/A1 semantic roles (~68k verb tokens)
- Participant tracking — 19 major figures tracked by subject/object/chapter
- Discourse structure — narrative peak scoring, episode boundaries (Longacre model)
- Derived stem morphology — Niphal, Piel, Pual, Hophal, Hithpael, Hiphil with full analysis suite
- Full queryable corpus (~480k words) with morphology, frequency, and concordance
- Per-book translation consistency analysis for Hebrew lemmas
- Word-level alignment between Hebrew OT and LXX
- Mood usage — subjunctive constructions, infinitive types, present vs. aorist imperatives
- Participle analysis — tense × voice profiles; genitive absolutes; perfect participles
- Discourse particles — καί/δέ/ὅτι/ἵνα/γάρ/οὖν/ἀλλά classification and function profiles
- Demonstratives — οὗτος vs. ἐκεῖνος; near/far comparison by genre
- Coreference and anaphora — pronoun referent chains via MACULA (~14,471 tokens)
- Syntactic role search — subject/object/verb relations via MACULA semantic links
- Louw-Nida domain search — semantic domain queries across the Greek NT
- Biblical Aramaic — Peal/Haphel/Pael verb stems; Daniel vs. Ezra morphological comparison
- Peshitta NT — 109,640 Syriac words with stem, tense/aspect, verbal morphology, and Sedra lemma
- Targumim — Onkelos, Targum Jonathan, and Targum to Psalms; 14,339 verses searchable by keyword
- Divine names and christological titles — OT/NT frequency with speaker and genre filters
- Intertextuality network — OT verse → NT citation bipartite graph
- Genre comparison — morphological patterns across narrative, epistle, prophecy, poetry
- Speech act classification (Searle taxonomy) — assertive/directive/commissive/expressive/declarative
- Information structure — parataxis/hypotaxis ratios, fronted elements, particle density (OT + NT)
- Stylometrics and register — TTR/MSTTR, wayyiqtol density, participle ratios, authorship fingerprinting
- Formulaic language — prophetic formulas, n-gram detection, blessing/curse patterns
For full API documentation and code examples, see docs/features.md.
| Source | Coverage |
|---|---|
| STEPBible TAHOT | ~284k Hebrew OT words with full morphology (stem, conjugation, PGN, state, Strong's) |
| STEPBible TAGNT | ~142k Greek NT words (tense/voice/mood/case/PGN) — Byzantine / Textus Receptus |
| STEPBible TALXX | ~480k LXX words with Strong's and Hebrew alignment |
| MACULA Hebrew | 475k words — WLC, syntax trees, semantic roles, LXX word alignment |
| MACULA Greek | 137k words — Nestle1904, syntax trees, semantic roles, referent links |
| Peshitta NT (ETCBC) | 109,640 Syriac words with full morphology |
| Targumim | Onkelos (Torah), Targum Jonathan (Prophets), Targum to Psalms — 14,339 verses |
| KJV | 24,570 English verses |
| Vulgate Clementine | 24,909 Latin verses |
Data is included as git submodules: stepbible-data/, macula-hebrew/, macula-greek/,
scrollmapper-data/, syrnt/. The processed database (data/processed/) is generated locally
by scripts/build_db.py and is not committed to the repo.
| Document | Contents |
|---|---|
| docs/data-sources.md | Data submodules, morphological coverage, data notes |
| docs/getting-started.md | Installation, prerequisites, first steps |
| docs/project-structure.md | Directory tree and module descriptions |
| docs/features.md | Full API reference — all 50+ analysis features with code examples |
| docs/lesson-packages.md | BBH / BBG / BBA lesson packages, exercise formats, PDF generation |
| notebooks/README.md | Jupyter notebook index by corpus and topic |
