A local tool that either combines a pile of Word / PowerPoint / Excel /
PDF / image files into one polished PDF, or lets you edit a single
existing PDF — rotate, crop, split, extract, reorder, redact, sign, stamp,
OCR, and compress. Everything runs on http://127.0.0.1:8765; no data leaves
your machine.
Derived from combinepdf. The original multi-file combine workflow still
works end to end; see CHANGES.md for what's new.
At startup a dialog asks what you want to do:
- Combine multiple files — the original workflow: pick files, set order, convert, reorder pages, then finish + export.
- Edit a single PDF — pick one PDF, then rotate/crop/split/extract/reorder, then finish + export. (Conversion is skipped entirely.)
Both modes share the optional Finishing steps: redaction, signatures, and stamps, followed by Settings (page numbers, bookmarks, metadata, OCR, compression, watermark, password).
- Python 3.9 or newer. (The project is pinned to stay 3.9-compatible.)
- Microsoft Office (Word/PowerPoint/Excel) for high-fidelity conversion in combine mode. Falls back to LibreOffice if found on PATH.
- Optional external binaries (see MIGRATION.md): Tesseract (OCR) and Ghostscript (aggressive compression + bordered-table fallback). Both are optional — features that need them degrade gracefully if absent.
Keep the project (and its venv) on a local disk, not a network drive.
Behind a corporate SSL-inspecting proxy, keep the --trusted-host flags.
cd path\to\enhanced-pdfer
py -3 -m venv .venv
.venv\Scripts\python -m pip install --upgrade pip --trusted-host pypi.org --trusted-host files.pythonhosted.org
.venv\Scripts\pip install -r requirements.txt --trusted-host pypi.org --trusted-host files.pythonhosted.org
.venv\Scripts\pip install -e . --no-deps
.venv\Scripts\python -m enhanced_pdferThe pip install -e . step registers the package (it lives under src/), which
is what makes python -m enhanced_pdfer resolve. --no-deps skips re-resolving
dependencies since requirements.txt already installed them.
If py -3 isn't available, use python or the full path to your interpreter.
cd path\to\enhanced-pdfer
.venv\Scripts\python -m enhanced_pdfercd path/to/enhanced-pdfer
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/pip install -e . --no-deps
.venv/bin/python -m enhanced_pdfer # subsequent runs: just this lineThere is no .bat / .command launcher — run the one-liners above directly.
<destination>/<Name>_<timestamp>/
├── Name.pdf ← final PDF
├── source_pdfs/ ← combine mode only: one PDF per source file
├── extracted.pdf ← single mode: "Extract selected" output
├── splits/ ← single mode: "Split" outputs
├── extracted_tables.xlsx ← single mode: "Extract tables" output
└── manifest.json ← record of inputs, settings, and final layout
- Soft-fail discipline. Every finishing/polish step (redaction, signatures, stamps, OCR, bookmarks, metadata, watermark, compression) is wrapped in try/except + a logged warning. The final PDF is the guaranteed deliverable. Password encryption is the one step that re-raises.
- PyMuPDF for page-level work. Rotation, crop, redaction, signatures, stamps, and page numbering all use PyMuPDF to avoid the QPDF segfault class of bugs seen when overlaying onto Office-converted pages.
- Real redaction. Redaction removes the underlying content from the
stream (PyMuPDF
apply_redactions), not just a black box on top. - Stamps vs watermark. Stamps are placed on specific pages at specific positions; the watermark is a full-deck diagonal overlay. Both exist.
- Timestamped logging. Every module logs through
_log.py; leave the cmd window open to watch live progress and errors.
enhanced-pdfer/
├── pyproject.toml
├── requirements.txt
├── README.md
├── CHANGES.md
├── MIGRATION.md
└── src/enhanced_pdfer/
├── __main__.py # entrypoint: mode dialog, pickers, server, browser
├── app.py # FastAPI app + stage-gated endpoints
├── session.py # in-memory session state + stage machine + manifest
├── pipeline.py # orchestrator (combine + single + finalize)
├── pickers.py # native OS dialogs (tkinter)
├── single_pdf.py # rotate / extract / split / crop / load (single mode)
├── redact.py # true redaction (rect + regex)
├── signature.py # signature image placement + type-to-sign
├── stamps.py # preset + custom stamps
├── ocr.py # ocrmypdf wrapper (graceful skip)
├── tables.py # pdfplumber + optional camelot -> xlsx
├── optimize.py # compression (standard + aggressive), watermark, password
├── merge.py # merge + bookmarks + metadata + page reorder
├── stamp.py # orientation-aware page numbers
├── blank_detect.py # visual-blank page scanner
├── excel_audit.py # print-area pre-flight
├── convert/ # platform conversion dispatch (COM-hardened on Win)
└── web/ # index.html, app.js, style.css