ੴ ਵਾਹਿਗੁਰੂ ਜੀ ਕੀ ਫ਼ਤਹਿ॥ ਸ੍ਰੀ ਭਗੌਤੀ ਜੀ ਸਹਾਇ॥
Converts the 4,300-page Fareedkot_Teeka.pdf from the legacy SriAngad font encoding to proper Unicode Gurmukhi, producing formatted .docx files for review. Also provides programmatic transliteration of the Teeka to Devanagari for reach to wider audiences.
Faridkot Wala Teeka (also spelled Faridkot Tika, Faridkoti Teeka, or Faridkot Vala Teeka; full title: Adi Sri Guru Granth Sahib ji Satik) is the earliest full-scale annotated commentary (exegesis or teeka) on the entire Sri Guru Granth Sahib. It provides detailed word-by-word explanations (arth), contextual backgrounds (uthankas), and interpretations of the Gurbani. It was produced in the late 19th and early 20th centuries under the patronage of the royal family of the princely state of Faridkot in Punjab, which is why it bears the name “Faridkot Wala” (“of Faridkot”).
It is widely regarded as a classical, traditional teeka produced by scholars of the Nirmala sect (a scholarly Sikh order known for deep engagement with Sanskrit, Braj, and Vedantic traditions). It served as the prototype and reference point for nearly all later commentaries on the Guru Granth Sahib.
The Guru Granth Sahib uses a mix of Punjabi, Hindi, and other languages in poetic form, with philosophical depth and linguistic features spanning centuries and regions. Earlier interpretations existed in oral form (mainly by Udasi and Nirmala scholars at gurdwaras) or in partial written works like glossaries (praydi), dictionaries (kos), Janam Sakhis, and Bhai Gurdas’s writings. The first notable written teekas appeared in the late 18th/early 19th century (e.g., Anandghana’s 1795 Japji teeka), but nothing covered the full scripture systematically.
The immediate trigger for the Faridkot Wala Teeka was the 1877 publication of a partial English translation of the Guru Granth Sahib by German scholar Ernest Trumpp. Many Sikhs viewed Trumpp’s work as disrespectful and inaccurate, with scorn toward Sikh interpretations and the faith itself. In response, Raja Bikram Singh (r. 1842–1898), ruler of Faridkot and a patron of the Amritsar Khalsa Diwan, commissioned a complete, authoritative Punjabi-language commentary to defend and clarify the scripture’s true meaning against colonial-era misinterpretations.
- Primary Drafter: Giani Badan Singh (also called Sant Giani Badan Singh Ji) of Sekhvari (or Dera Sekhwan). A Nirmala scholar, he prepared the first complete draft in about 6½ years; it was ready by 1883.
- Revision Committee: A synod of scholars representing diverse Sikh schools of thought reviewed and refined the draft. Mahant Sumer Singh (or Shamer Singh) of Patna Sahib served as chairman. Other members included:
- Giani Harbhajan Singh of Amritsar
- Sant Singh of Kapurthala state
- Jhanda Singh of Gurdwara Nanakiana Sahib (near Sangrur)
- Rai Singh of Jarigi Rana (or Jangi Rana)
- Dhian Singh of Sekhvari
- Pandit Hamir Singh Sariskriti
- Pandit Balak Ram Udasi Sariskriti
- Baba Bakhtavar Singh Giani
The revision process was completed during Raja Bikram Singh’s lifetime, though he did not live to see the printed edition.
Note: A contemporary scholar, Pandit Tara Singh Narotam, began a parallel teeka but died after completing only up to Basant Rag; it was never widely circulated.
The teeka is written in Braj Bhasha (an aristocratic literary language of the time, using Gurmukhi script), not simple modern Punjabi. This reflects the Nirmala scholarly tradition, which often drew on Sanskrit and classical Indic styles. While the Guru Granth Sahib itself is largely accessible to Punjabi speakers, the teeka’s Braj style made it more suited to educated elites and traditional scholars at the time. It includes detailed arth (meanings), uthankas (historical/spiritual contexts), and explanations that sometimes incorporate Vedantic or Upanishadic philosophical frameworks common in Nirmala exegesis.
- First Edition (early 1900s): Funded by Raja Balbir Singh (successor to Bikram Singh). Printed in four volumes at Wazir Hind Press, Amritsar (founded by Bhai Vir Singh). Three volumes appeared during Balbir Singh’s reign; the fourth under Maharaja Brijindar Singh. Many copies were distributed free to gurdwaras and scholars; the rest sold at nominal cost. The full set is approximately 4,300 pages.
- Second Edition (1924/1928): Published by Maharaja Harinder Singh. Singh Sabha reformers had suggested revisions (e.g., switching to standard modern Punjabi), and a committee was briefly formed in 1918, but the Maharaja insisted on preserving the original form and style. A second edition was released without major changes.
- Later Reprints: The original manuscript is preserved in the Faridkot royal toshakhana. The Languages Department (Bhasha Vibhag), Punjab, published reprints starting in 1970. Modern printed sets (usually 4 volumes) are still available from publishers like Singh Brothers.
Digital versions (scans of the 1924 print and others) are freely available on Archive.org and other Sikh resource sites.
The teeka covers the entire Guru Granth Sahib page-by-page. For each shabad or pauri:
- It gives the literal meaning.
- Provides historical or spiritual context (uthanka).
- Explains philosophical and doctrinal nuances.
- Draws on traditional Nirmala scholarship while staying rooted in Gurbani.
It is praised for its systematic approach, separating interpretive commentary from historical notes, and for being the first comprehensive written reference of its kind.
- It is the first complete teeka and became the model (“ideal prototype”) for all subsequent commentaries.
- It remains highly respected in traditional Sikh sampardas (Nirmala, Taksal, Nanaksar, etc.) and is frequently cited in katha (expositions) by scholars like Bhai Seva Singh.
- It helped preserve and standardize traditional interpretations during a time of colonial influence and reform movements.
- Many later teekas (including those by Prof. Sahib Singh or others) built upon or reacted to it.
Because of its Nirmala authorship and use of Braj with occasional Vedantic/Brahmanical philosophical undertones, it faced criticism from Singh Sabha reformers and some modern Sikh scholars who preferred a more “purely Sikh” or literal interpretation without external influences. Critics argued it made the Gurbani seem overly complex or aligned too closely with Hindu philosophical frameworks. Despite this, it is still valued as a classical resource, especially for its detailed uthankas and traditional insights. Some contemporary voices (e.g., in forums or katha) defend it strongly as the “finest” or most authoritative traditional commentary.
- Print: 4-volume sets (e.g., from Singh Brothers or Khalsa Shop).
- Digital: Free PDFs on Archive.org (including the 1924 edition and scanned originals), Vidhia.com, and Gurmat Veechar.
- Audio/Katha: Many modern kathavachaks (e.g., from Nanaksar samparda) base explanations on it; recordings are available online.
In short, the Faridkot Wala Teeka stands as a monumental scholarly achievement in Sikh exegetical literature — a direct response to external challenges, a product of royal patronage and collaborative Nirmala scholarship, and a foundational text that continues to shape traditional understanding of the Guru Granth Sahib more than 140 years after its creation. If you want links to specific volumes, excerpts, or comparisons with other teekas, let me know!
For generations, this work has been accessed only in its original printed form — a rare and fragile manuscript available in limited copies. The digitization and modernization of this text is critical for:
- Preservation — Protect this irreplaceable scholarly work from deterioration and loss
- Accessibility — Make the Teeka available to Sikhs worldwide, especially younger generations
- Searchability — Enable digital search, study, and reference capabilities
- Multilingual Reach — Provide Hindi/Devanagari versions for Hindi-speaking communities who can benefit from this wisdom
The PDF uses SriAngad — a legacy pre-Unicode Gurmukhi font (from pre-2000) that maps Gurmukhi glyphs to Latin character codes (WinAnsiEncoding). It has no ToUnicode map, so:
- Copy-paste from the PDF yields Latin gibberish (e.g.,
manmuKhinstead of ਮਨਮੁਖ) - The text cannot be indexed, searched, or properly displayed on modern devices
- Standard OCR and digitization tools fail completely
This project reverse-engineers the complete SriAngad encoding and converts the text to proper Unicode, enabling the Teeka to be used with modern software and devices.
See ANALYSIS.md for the full character mapping reference.
Generated files (output/*.docx) are programmatically converted from PDF and may contain:
- Spelling mistakes or clerical errors from the source document
- Character mapping errors (rare, but possible edge cases)
- Formatting inconsistencies where the PDF structure was ambiguous
A comprehensive manual proofreading process is planned to verify and correct all text. This conversion is a first pass to enable digital distribution and review — it is NOT a final, verified edition.
For questions about specific passages or suspected errors, please refer to the original PDF at the noted page number.
Found an error? Please report it:
- Email: sksingh2211@gmail.com (include page number and description of the error)
- GitHub: Open an issue at https://github.com/qascade/faridkot-teeka/issues
Your feedback helps improve the accuracy of the conversion for future readers.
The complete Fareedkot Teeka is now available in Unicode format:
- fareedkot_teeka_0_1_0.docx — Full 4,300-page document in Gurmukhi Unicode with formatting
- fareedkot_teeka_0.1.0.pdf — PDF version for easier sharing
- fareedkot_teeka_devanagari_0_1_0.docx — Complete transliteration to Devanagari Unicode (v0.1.0)
- fareedkot_teeka_devanagari_0.1.0.pdf — PDF version for wider compatibility
fareedkot_teeka_devanagari_0_2_0.docx(in progress) — v0.2.0 with Devanagari spelling corrections viabraj_correction_agent/
All documents are annotated with version metadata and include proper formatting with page references to the original GGS.
# Create virtual environment
python3 -m venv venv
# Activate it
source venv/bin/activate # On macOS/Linux
# or
venv\Scripts\activate # On Windowspip install -r requirements.txtRequirements: pdfminer.six, python-docx
deactivateFont: Uses Noto Sans Gurmukhi (cross-platform). On other systems, edit _GURMUKHI_FONT in src/generate.py.
Run from the project root (so the PDF and output/ folder are found correctly):
# 10 pages starting at page 338
python3 src/generate.py -s 338 -e 347
# Custom output path
python3 src/generate.py -s 1 -e 50 --output output/first_50.docx
# With page breaks between pages
python3 src/generate.py -s 338 -e 347 --page-break
# Only Gurbani and Braj (skip refs/annotations)
python3 src/generate.py -s 338 -e 347 --types gurbani braj
# No page number separators
python3 src/generate.py -s 338 -e 347 --no-page-markers
# Quiet (no per-page progress)
python3 src/generate.py -s 338 -e 347 -qOutput goes to output/pages_{start}-{end}.docx by default.
Convert an existing Gurmukhi .docx file to Devanagari script (for Hindi-reading audiences):
python3 src/transliterate_docx.py input.docx [output.docx]This auto-detects Gurmukhi text, transliterates to Devanagari Unicode, and changes the font to Noto Serif Devanagari. Language tag is updated from pa-IN to hi-IN.
Example:
python3 src/transliterate_docx.py output/pages_338-347.docx
# → Saves to: output/pages_338-347_devanagari.docxThe braj_correction_agent/ folder contains a LangGraph + Gemini pipeline that incrementally
fixes spelling errors in the Devanagari DOCX. It processes the Braj commentary paragraphs in
batches, pausing after each one for human review before writing any changes. Progress is saved
to a SQLite checkpoint so runs can be stopped and resumed freely.
pip install -r braj_correction_agent/requirements.txt
# Add your Gemini API key to braj_correction_agent/.env, then:
python braj_correction_agent/main.py \
--input "Finished Docs/fareedkot_teeka_devanagari_0_1_0.docx"See braj_correction_agent/README.md for full usage.
Compare every Gurbani line in a generated docx against the authoritative Unicode GGS from GurbaniDB:
# One-time: download all 1430 GGS pages as reference (~15 seconds)
python3 gurbani_accuracy_test/download_ggs_reference.py
# Run accuracy test
python3 gurbani_accuracy_test/test_gurbani_accuracy.py "Finished Docs/fareedkot_teeka_0_1_0.docx"Results are saved as JSON + markdown report in gurbani_accuracy_test/data/. See docs/gurbani_accuracy_report.md for the latest report.
Current accuracy: 98.9% across 60,726 tuks tested (v0.1.2). Target for v0.2.0: spelling corrections via braj_correction_agent/.
python3 src/test_converter.py # 171 character mapping tests
python3 src/test_transliterator.py # 81 transliteration testsAll tests must pass before any changes to src/converter.py or src/transliterator.py.
| File | Purpose |
|---|---|
src/converter.py |
Pure SriAngad → Unicode conversion (no PDF assumptions). Can be used as a standalone API. |
src/extractor.py |
Extracts text + color/font metadata from PDF using pdfminer. Cleans PDF artifacts via normalize_pdf_text() before conversion. |
src/generate.py |
CLI — main entry point for generating .docx files |
src/docx_generator.py |
Renders Element objects into a .docx |
src/transliterator.py |
Gurmukhi Unicode → Devanagari Unicode transliteration engine |
src/transliterate_docx.py |
CLI for batch Word document transliteration |
gurbani_accuracy_test/ |
Accuracy test framework — compares docx gurbani against GurbaniDB reference |
src/test_converter.py |
171 unit tests covering every character mapping |
src/test_transliterator.py |
81 unit tests for transliteration (all passing) |
requirements.txt |
Python dependencies |
docs/ANALYSIS.md |
Complete SriAngad → Unicode character mapping reference |
docs/doc_generation_instructions.md |
Page structure rules + known common errors |
docs/sample_mappings.md |
Ground-truth Gurbani lines used to verify mappings |
docs/ |
Reference PDFs (font key maps) |
Text is classified by color (from PDF graphicstate) and font size:
| Type | Color in PDF | Style in output |
|---|---|---|
gurbani |
Deep blue (0, 0, 0.502) |
Blue bold centered, 16pt |
braj |
Red (0.502, 0, 0) or black ≥10pt |
Dark red justified, 12pt |
braj_sub |
Bright blue (0, 0, 1.0) |
Blue justified, 11pt |
ggs_ref |
Magenta (1, 0, 1) |
Pink, 11pt |
annotation |
Black < 10pt | Gray, 9pt |
latin |
Times New Roman font | Black, 10pt |
- Sihari reordering —
iis coded BEFORE its consonant in SriAngad; must be placed AFTER in Unicode. With subscript clusters:i+ base + subscript → base + subscript + ਿ - Carrier + matra look-ahead —
a/e/Afollowed by a matra produce precomposed independent vowels (e.g.aw→ਆ,ey→ਏ,Au→ਉ) - Ø+z sequence — together produce ੱ (addak); Ø alone has no output
- Matra ordering — tippi/bindi typed before dulainkar/aunkar, and aunkar typed before hora/kanaura, are swapped to correct Unicode order
Separation of Concerns:
converter.py— Pure SriAngad → Unicode mapping. No PDF-specific logic. Can be used as a standalone API.extractor.py— PDF extraction + artifact cleanup vianormalize_pdf_text()→ passes clean text toconvert().
This design lets users leverage converter.py directly for their own SriAngad sources without PDF assumptions.
Data flow:
PDF text → normalize_pdf_text() → convert() → Unicode Gurmukhi
- Page markers (
─── Page N ───) are included for reference; remove after verification - Hindi/Devanagari versions — The transliteration is direct character-by-character conversion, not semantic translation. This produces valid Devanagari text but with non-standard spellings in some places (matra ordering, anusvara, nukta artifacts). The
braj_correction_agent/pipeline is being used to fix these incrementally with human review. - Subscript Ha (੍ਹ) — ~113 instances where the subscript Ha may be lost during PDF extraction. Under investigation.
(8.0, )