Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🏛️ Latin Dictionary

A dictionary combining the complete Lewis & Short Latin Dictionary (51,636 unabridged entries) with a Synonyms & Near-Synonyms section, always-visible morphology tables, mined usage/register labels, and 388 grammar reference entries from Allen & Greenough's New Latin Grammar.

It ships in two forms, from the same source data:

  • macOS — a .dictionary plugin for the native Dictionary app and the system-wide "Look Up" feature.
  • Linux & Windows — a StarDict build for GoldenDict-ng, whose Scan Popup gives the same select-a-word-and-look-it-up workflow. See docs/GOLDENDICT.md.

v1.7.0 — Linux and Windows support via StarDict/GoldenDict-ng, built from the same XML as the macOS bundle: all 54,570 entries and all 410,315 Morpheus-indexed inflected forms, so selecting puellarum in a text resolves to puella exactly as it does on macOS. Allen & Greenough cross-references are rewritten from Apple's x-dictionary: scheme to bword:, and the stylesheet is re-cut for GoldenDict with a dark-mode variant, cross-platform font stacks, and collapsed sense indentation at popup width. The installed footprint drops from ~116 MB to 22 MB, since the article body ships dictzip-compressed. The project also dropped its -mac suffix, which stopped being accurate with this release; GitHub redirects the old URLs.

v1.6.0 — Adds a Pronunciation Guide reference entry: five tables (Vowels, Nasalized Vowels, Diphthongs, Consonants, Geminates) mapping this dictionary's own reconstructed IPA onto real example words in English, German, French, Italian, Spanish, and Icelandic wherever a genuinely close sound exists, plus a 7th Ecclesiastical Latin column on rows where Church pronunciation diverges instructively from Classical. Geminates are checked against this project's own 51,636-headword corpus rather than assumed. Also turns Allen & Greenough's own <ref n="NNN"> cross-references — 1,144 of them, previously rendered as plain unclickable text — into real in-dictionary links wherever the target is one of this project's own indexed grammar entries (172 qualify; the rest correctly stay plain text rather than becoming dead links).

v1.5.0 — Reconstructed Classical Latin pronunciation, rule-based per W. Sidney Allen's Vox Latina and checked directly against Wiktionary's Module:la-pronunc output: vowel length/quality, c/k/qu → [k], v → [w], dark vs. clear L (lupus → [ˈɫʊ.pʊs], insula → [ˈĩː.sʊ.ɫa]), nasalized vowels before final -m or before n/m+s/f (etiam → [ɛ.ti.ãː]), hiatus-vowel tensing (duo → [ˈdu.ɔ]), and the classical penultimate-weight stress rule — shown for the ~85% of entries carrying at least one macron/breve, skipped rather than guessed elsewhere. Also fixes two silent data gaps: nouns whose Morpheus lemma carries a homonym-disambiguating number absent from L&S's own key (equusequus1) were missing their declension table entirely; semi-deponent verbs (audeo, gaudeo, soleo, fido and compounds) now get their correct 3-part citation (audeo, audere, ausus sum) instead of a broken one, since L&S's own tagging never marks this class as deponent.

v1.4.1 — Adjective comparison (positive/comparative/superlative) is now properly modeled: the positive entry shows a one-line degree summary (bonus · melior · optimus) and its own gender-split declension table, while melior and optimus each get their own focused synthetic entry with a full declension table and search index — including the five suppletive irregulars (bonus/melior/optimus, malus/peior/pessimus, magnus/maior/maximus, parvus/minor/minimus, multus/plus/plurimus) that Perseus/Morpheus doesn't tag consistently. Also fixes a search bug where a headword's own displayed spelling with macrons/breves (bŏnus) wasn't indexed — only the marks-stripped form (bonus) was searchable before.

v1.4.0 — Adds Döderlein's Hand-book of Latin Synonymes as a third synonym source (549 articles), subjunctive/imperative/gerundive morphology, quoted-example citations split into distinct Latin quote / English gloss / reference, real HTML paradigm tables in the grammar entries (previously garbled run-on text), a Ramshorn "Latin Terminations" word-formation reference, and styled cross-reference markers.

✨ Features

  • 51k Unabridged Lewis & Short Entries: The complete 1879 Harpers'/Oxford A Latin Dictionary compiled from Perseus TEI-XML into the macOS .dictionary format.
  • Synonyms & Near-Synonyms: three complementary sources attached to the relevant entries — a compact near-synonym list (with declension/conjugation markup) from Spinelli–Fenzi's First Online Dictionary of Latin Near-Synonyms (St Andrews, 2019), 1,015 discussion articles from Ramshorn's Dictionary of Latin Synonymes (tr. Lieber, 1841), and 549 discussion articles from Döderlein's Hand-book of Latin Synonymes (tr. Arnold, 1874) — e.g. amo shows both Ramshorn's and Döderlein's independent treatments of Amare/Diligere/Caritas distinguishing the shades of meaning.
  • Principal Parts: Verbs lead with the classic citation form Latin is taught with — amo, amare, amavi, amatus — or for deponents, sequor, sequi, secutus sum. Deponent detection comes from L&S's own part-of-speech tag; the 4th part falls back to the future active participle for intransitive verbs with no perfect passive participle.
  • Morphology Tables: Nouns and adjectives get a Case × Number declension grid (split into separate Masculine/Feminine/Neuter tables when more than one gender is attested, e.g. adjectives — plain single-gender nouns still get one flat table); verbs get an indicative (1st sg.) table across all six tenses in both voices, plus infinitives, subjunctive (1st sg., 4 tenses), imperative (2nd sg./pl., present & future), and the gerundive — built from the Perseus/Morpheus full-form analyses (392k forms). Subjunctive/imperative are shown only where actually attested, since Morpheus's corpus-attested coverage is inconsistent across verbs.
  • Adjective Comparison (Positive/Comparative/Superlative): the positive-degree entry shows a compact summary line (Positive: bonus · Comparative: melior · Superlative: optimus) plus its own declension table; the comparative and superlative each get their own focused synthetic entry — own headword, own gender-split declension table, own search index — rather than three tables crammed onto one page. Covers both regular formation (altus → altior → altissimus) and the five suppletive irregulars whose superlative Morpheus doesn't tag consistently (bonus/melior/optimus, malus/peior/pessimus, magnus/maior/maximus, parvus/minor/minimus, multus/plus/plurimus).
  • Reconstructed Pronunciation: a Vox Latina-style IPA transcription appears under the headword for the ~85% of entries with at least one macron/breve marked (skipped rather than guessed for the rest) — vowel length and quality, dark vs. clear L, nasalized vowels before final -m/before n-m+s-f, hiatus-vowel tensing, and classical stress placement, checked directly against Wiktionary's own generated pronunciations rather than derived from rules alone. A Pronunciation Guide reference entry maps every vowel, diphthong, consonant, and geminate onto real example words in English, German, French, Italian, Spanish, and Icelandic (plus Ecclesiastical Latin where it instructively diverges from Classical).
  • Usage & Register Labels: L&S's own <usg> markup is mined and styled instead of rendered as flat text — technical-domain labels (Military/Medical/Mercantile/Political term) surface as badges under the headword; rhetorical labels (Lit./Transf./Trop./Poet./Meton.) and inline case/mood/number abbreviations get distinct styling inline.
  • Etymology: L&S's <etym> markup — the source root or foreign-language origin (e.g. abactor < [abigo]; Abaddir < [Heb. אָב אַדּיִרּ, mighty father]) — is styled and bracketed in the preamble instead of appearing as unmarked plain text.
  • Quoted Examples: every <cit> (a real Latin example quote, its English gloss, and the citing author/work) renders as three distinct pieces — italicized Latin, curly-quoted translation, small grey reference — instead of one undifferentiated blob.
  • Cross-References: L&S's <xr>/<ref> markers (e.g. "v. supra/infra") are styled instead of rendered as flat text.
  • Grammar Reference (Allen & Greenough): 388 entries from A New Latin Grammar for Schools and Colleges (1903) — both broad topics (subsections/subsubsections, e.g. "The Locative Case") and precise named rules (e.g. "Ablative Absolute", "Hortatory Subjunctive", "Historical Present"), each look-up-able by its topic name or by citation ("AG 419", "§419") — the way commentaries actually reference it. Includes real HTML paradigm tables (noun declensions, sound classifications), not flattened text. A&G's own internal cross-references (<ref n="419">) are clickable links wherever the target is itself an indexed entry, not just styled text.
  • Word-Formation Reference (Ramshorn Terminations): 20 entries from the front matter of Ramshorn's Dictionary of Latin Synonymes — a guide to what Latin suffixes mean (e.g. -tas designates quality, -tor designates an agent), look-up-able as "Latin Terminations I" through "XXIV".
  • Inflected-Form Lookup: Every attested inflected form is indexed, so macOS "Look Up" (Force Click / Three-Finger Tap) works from any word in a real Latin text, not just dictionary headwords. Orthographic variants (i/j, u/v) are indexed too.
  • System Integration: Works natively with Dictionary.app and system-wide Look Up.

📦 Installation (For End Users)

  1. Download the latest release from the Releases page.
  2. Unzip and drag LatinDictionary.dictionary into ~/Library/Dictionaries/.
  3. Open the macOS Dictionary app, go to Settings, and enable "Latin (Lewis & Short)".

🐧 🪟 Linux and Windows

The same data ships as StarDict for GoldenDict-ng, which runs on Linux, Windows and macOS — all 54,570 entries and all 410,315 indexed inflected forms. GoldenDict's Scan Popup is the direct analogue of macOS Look Up: select a word anywhere on screen and get the entry, so reading a real Latin text works the same way it does on macOS.

Download Latin-GoldenDict-<version>.zip from the Releases page. docs/GOLDENDICT.md has the full setup — adding the folder, installing the stylesheet (StarDict has no stylesheet slot, so it ships separately), enabling Scan Popup, and the Wayland caveat.

🛠️ Building from Source

Prerequisites

  • Python 3.x (standard library only — no packages needed)
  • Dictionary Development Kit (in Apple's "Additional Tools for Xcode"), expected at /Applications/XcodeAdditionalTools/Utilities/DictionaryDevelopmentKit (rename /Applications/Additional Tools for Xcode — no spaces in the path)
  • macOS with Xcode command-line tools

Step 1 — Download the source data

None of this is committed to the repo (see .gitignore):

  1. Lewis & Short TEI-XML (CC BY-SA 4.0) — download into data/lewis_short/:

    mkdir -p data/lewis_short
    curl -L -o data/lewis_short/lat.ls.perseus-eng2.xml \
      https://raw.githubusercontent.com/PerseusDL/lexica/master/CTS_XML_TEI/perseus/pdllex/lat/ls/lat.ls.perseus-eng2.xml
  2. Morphology (Perseus/Morpheus full-form analyses, via the Diogenes prebuilt data) — extract latin-lemmata.txt into data/analyses/:

    curl -L -o /tmp/prebuilt-data.tar.xz \
      https://github.com/pjheslin/diogenes-prebuilt-data/raw/master/prebuilt-data.tar.xz
    mkdir -p data/analyses
    tar -xf /tmp/prebuilt-data.tar.xz -C /tmp dependencies/data/latin-lemmata.txt
    mv /tmp/dependencies/data/latin-lemmata.txt data/analyses/
  3. Ramshorn synonyms (public domain) — OCR text from archive.org into data/ramshorn/:

    mkdir -p data/ramshorn
    curl -L -o data/ramshorn/ramshorn_1841_djvu.txt \
      "https://archive.org/download/ramshorn-lewis-dictionary-of-latin-synonymes-1841/RAMSHORN%2C%20Lewis%20-%20Dictionary%20Of%20Latin%20Synonymes%20%5B1841%5D_djvu.txt"
  4. Spinelli–Fenzi near-synonyms (CC BY) — already committed at data/spinelli/latin_near_synonyms.json (86 KB); nothing to download.

  5. Allen & Greenough grammar (CC BY-SA 4.0) — the original Perseus TEI-XML (not the later NC-licensed Alpheios edition; see LICENSE) into data/allen_greenough/:

    mkdir -p data/allen_greenough
    curl -L -o data/allen_greenough/ag_grammar.xml \
      https://raw.githubusercontent.com/PerseusDL/canonical-pdlrefwk/master/data/viaf39744457/001/viaf39744457.001.perseus-eng1.xml
  6. Döderlein's Hand-book of Latin Synonymes (public domain) — Project Gutenberg EBook #33197 into data/doederlein/:

    mkdir -p data/doederlein
    curl -L -o data/doederlein/doederlein_gutenberg.txt \
      https://www.gutenberg.org/cache/epub/33197/pg33197.txt

Step 2 — Build the databases and dictionary XML

python3 scripts/build_dbs.py      # L&S / analyses / Ramshorn / Spinelli / Doederlein -> SQLite (ls.db, morph.db, synonyms.db)
python3 scripts/build_grammar.py  # A&G TEI-XML -> data/grammar.db
python3 scripts/build_xml.py      # SQLite -> src/LatinDictionary.xml (~106 MB)
python3 scripts/test_dictionary.py  # regression checks (XML validity + known-entry assertions)

Step 3 — Compile and install

cd src
make install

Then open Dictionary.app → Settings and enable "Latin (Lewis & Short)".

(If the DDK errors with unable to parse objects/dict.plist, simply re-run make — it is a transient failure of the kit's xsltproc step.)

📁 Project Structure

latin-dictionary/
├── data/
│   ├── lewis_short/           # Perseus L&S TEI-XML [gitignored]
│   ├── analyses/              # Morpheus full-form data via Diogenes [gitignored]
│   ├── ramshorn/              # Ramshorn 1841 OCR text [gitignored]
│   ├── spinelli/              # Spinelli-Fenzi near-synonyms JSON (committed, 86 KB)
│   ├── doederlein/            # Doederlein 1874 Gutenberg text [gitignored]
│   ├── allen_greenough/       # A&G Perseus TEI-XML [gitignored]
│   ├── morpheus/              # Morpheus engine checkout, optional [gitignored]
│   ├── ls.db                  # SQLite L&S entries [gitignored]
│   ├── morph.db               # SQLite full-form morphology [gitignored]
│   ├── synonyms.db            # SQLite Ramshorn + Spinelli + Doederlein articles [gitignored]
│   └── grammar.db             # SQLite A&G sections + named rules [gitignored]
├── scripts/
│   ├── build_dbs.py           # L&S / analyses / Ramshorn / Spinelli / Doederlein -> SQLite
│   ├── build_grammar.py       # A&G TEI-XML -> grammar.db
│   ├── build_xml.py           # SQLite -> Apple Dictionary XML
│   ├── build_stardict.py      # Apple XML -> StarDict (Linux/Windows)
│   ├── verify_stardict.py     # Independent reader for the built StarDict set
│   ├── dictzip.py             # Random-access gzip for the StarDict body
│   ├── package_goldendict.sh  # One-command GoldenDict release bundle
│   ├── test_dictionary.py     # Regression checks on the generated XML
│   └── install_dictionary.sh  # One-command build & install
├── src/
│   ├── LatinDictionary.xml    # Generated dictionary source [gitignored]
│   ├── LatinDictionary.css    # Dictionary styling (macOS)
│   ├── GoldenDictArticle.css  # Dictionary styling (GoldenDict)
│   ├── LatinDictionary.plist  # Apple Dictionary metadata
│   ├── Makefile               # Build rules
│   └── objects/               # Build artifacts [gitignored]
├── docs/
│   └── GOLDENDICT.md          # Linux/Windows install guide
└── README.md

📚 Data Sources

  • Lewis & Short: A Latin Dictionary (1879), TEI-XML from the Perseus Digital Library — CC BY-SA 4.0, with funding from The National Endowment for the Humanities.
  • Morphology: Full-form analyses generated by the Perseus Morpheus analyzer, as packaged by Diogenes.
  • Synonyms & Word-Formation: Ramshorn, Dictionary of Latin Synonymes, for the use of schools and private students, translated by Francis Lieber (Boston, 1841) — public domain, OCR text from the Internet Archive. Its front-matter "Terminations" guide (word-formation suffixes) is parsed alongside the synonym articles.
  • Near-Synonyms: Spinelli & Fenzi, The First Online Dictionary of Latin Near-Synonyms (University of St Andrews, 2019) — CC BY, DOI 10.17630/3cf644e6-86b8-44d0-a50a-b33c7ca86072.
  • Synonyms (second source): Döderlein, Hand-book of Latin Synonymes, translated by H. H. Arnold (1874) — public domain, Project Gutenberg EBook #33197. An independent 19th-century treatment alongside Ramshorn's, often organized and phrased differently for the same headwords.
  • Grammar: Allen, Greenough, Kittredge, Howard & D'Ooge, A New Latin Grammar for Schools and Colleges (Ginn & Co., 1903) — the original Perseus Digital Library TEI-XML digitization, CC BY-SA 4.0. (Not the later Dickinson College Commentaries / Alpheios revision, which is CC BY-NC-SA and therefore not redistributable here.)

🤝 Contributing

Contributions are welcome. The most valuable are:

  • Weird/broken entries: With 51,636 entries auto-generated from TEI-XML and an 1841 OCR text, edge cases slip through (mangled preambles, mis-parsed synonym articles, garbled OCR in the synonyms sections). Open an issue with the headword and a screenshot, or trace it to scripts/build_xml.py / scripts/build_dbs.py and send a PR.
  • Ramshorn OCR cleanup: The synonym article text comes from OCR and retains scanning errors. Corrections to the parsing heuristics in build_dbs.py (or a cleaner source text) would improve every affected entry. (Döderlein comes from a clean Project Gutenberg transcription, so this doesn't apply there.)
  • Styling: Enhance src/LatinDictionary.css.

📄 License

  • Code (Python scripts, CSS, Makefile): MIT License
  • Lewis & Short text: CC BY-SA 4.0 (per Perseus; see their availability statement)
  • Ramshorn synonyms: public domain
  • Döderlein synonyms: public domain (Project Gutenberg)
  • Morphology data: Perseus/Morpheus, CC BY-SA 4.0
  • Allen & Greenough grammar: Perseus TEI-XML edition, CC BY-SA 4.0

When distributing this dictionary, the Perseus attribution must remain intact:

Text provided under a CC BY-SA license by Perseus Digital Library, http://www.perseus.tufts.edu, with funding from The National Endowment for the Humanities. Data accessed from https://github.com/PerseusDL/lexica/.

🙏 Acknowledgments

  • Charlton T. Lewis & Charles Short, and the Perseus Digital Library for the digitized text.
  • Ludwig Ramshorn & Francis Lieber for the synonyms handbook; the Internet Archive for the scan.
  • Ludwig Döderlein & H. H. Arnold for the second synonyms handbook; Project Gutenberg for the clean transcription.
  • Peter Heslin (Diogenes) for the prebuilt Morpheus analyses.
  • Tommaso Spinelli & Giacomo Fenzi for the near-synonyms dataset.
  • J.B. Greenough, G.L. Kittredge, A.A. Howard & Benjamin L. D'Ooge, editors of Allen & Greenough's grammar.
  • Apple Dictionary Development Kit for the .dictionary format tooling.

About

Complete Lewis & Short Latin lexicon with Morpheus morphology, Ramshorn/Spinelli synonyms and reconstructed pronunciation, for the macOS Dictionary app and system-wide Look Up — select any inflected form and look it up.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages