PageLedger 0.5.2 records extraction attempts, source identity, cost, and review work as plain files. Choose an entry point below.
| Task | Guide |
|---|---|
| Try PageLedger without an OCR engine | First run: text, review, rerun, and replay |
| Run a complete offline document job | First document job: process, recover, inspect, verify, and review |
| Extract a PDF text layer or scanned PDF | First PDF/OCR run |
| Process a document through local text, OCR, and optional image stages | Document processing jobs |
| Recover interrupted extraction in place | Checkpoint recovery |
| Review selected text and record human decisions | Document reports and review receipts |
- Choose an OCR or VLM adapter
- Classify pages and review route evidence
- Run OCR on non-English and historical documents
- Work through a scanned government archive
- Write a custom extraction adapter
- Compare PageLedger with extraction tools
- CLI commands and configuration
- Run directory and artifact overview
- Run manifest
- Route map
- Provenance and quality JSONL
- Normalized records
- Audit queue
- Rerun manifest
- Checkpoint and page-attempt records
- Document job
- Document report and human review receipts
- Reader trial protocol and results form
- Page image input evidence
The JSON Schemas define the machine-readable artifact contract.
Package 0.5.2 retains artifact schema_version: "0.1".