Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

enhanced-pdfer

A local tool that either combines a pile of Word / PowerPoint / Excel / PDF / image files into one polished PDF, or lets you edit a single existing PDF — rotate, crop, split, extract, reorder, redact, sign, stamp, OCR, and compress. Everything runs on http://127.0.0.1:8765; no data leaves your machine.

Derived from combinepdf. The original multi-file combine workflow still works end to end; see CHANGES.md for what's new.

Two modes

At startup a dialog asks what you want to do:

  • Combine multiple files — the original workflow: pick files, set order, convert, reorder pages, then finish + export.
  • Edit a single PDF — pick one PDF, then rotate/crop/split/extract/reorder, then finish + export. (Conversion is skipped entirely.)

Both modes share the optional Finishing steps: redaction, signatures, and stamps, followed by Settings (page numbers, bookmarks, metadata, OCR, compression, watermark, password).

Requirements

  • Python 3.9 or newer. (The project is pinned to stay 3.9-compatible.)
  • Microsoft Office (Word/PowerPoint/Excel) for high-fidelity conversion in combine mode. Falls back to LibreOffice if found on PATH.
  • Optional external binaries (see MIGRATION.md): Tesseract (OCR) and Ghostscript (aggressive compression + bordered-table fallback). Both are optional — features that need them degrade gracefully if absent.

First run (Windows, cmd)

Keep the project (and its venv) on a local disk, not a network drive. Behind a corporate SSL-inspecting proxy, keep the --trusted-host flags.

cd path\to\enhanced-pdfer
py -3 -m venv .venv
.venv\Scripts\python -m pip install --upgrade pip --trusted-host pypi.org --trusted-host files.pythonhosted.org
.venv\Scripts\pip install -r requirements.txt --trusted-host pypi.org --trusted-host files.pythonhosted.org
.venv\Scripts\pip install -e . --no-deps
.venv\Scripts\python -m enhanced_pdfer

The pip install -e . step registers the package (it lives under src/), which is what makes python -m enhanced_pdfer resolve. --no-deps skips re-resolving dependencies since requirements.txt already installed them.

If py -3 isn't available, use python or the full path to your interpreter.

Subsequent runs (Windows, cmd)

cd path\to\enhanced-pdfer
.venv\Scripts\python -m enhanced_pdfer

First run / subsequent runs (macOS, nice-to-have)

cd path/to/enhanced-pdfer
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/pip install -e . --no-deps
.venv/bin/python -m enhanced_pdfer       # subsequent runs: just this line

There is no .bat / .command launcher — run the one-liners above directly.

Output layout

<destination>/<Name>_<timestamp>/
├── Name.pdf                  ← final PDF
├── source_pdfs/              ← combine mode only: one PDF per source file
├── extracted.pdf             ← single mode: "Extract selected" output
├── splits/                   ← single mode: "Split" outputs
├── extracted_tables.xlsx     ← single mode: "Extract tables" output
└── manifest.json             ← record of inputs, settings, and final layout

Design notes

  • Soft-fail discipline. Every finishing/polish step (redaction, signatures, stamps, OCR, bookmarks, metadata, watermark, compression) is wrapped in try/except + a logged warning. The final PDF is the guaranteed deliverable. Password encryption is the one step that re-raises.
  • PyMuPDF for page-level work. Rotation, crop, redaction, signatures, stamps, and page numbering all use PyMuPDF to avoid the QPDF segfault class of bugs seen when overlaying onto Office-converted pages.
  • Real redaction. Redaction removes the underlying content from the stream (PyMuPDF apply_redactions), not just a black box on top.
  • Stamps vs watermark. Stamps are placed on specific pages at specific positions; the watermark is a full-deck diagonal overlay. Both exist.
  • Timestamped logging. Every module logs through _log.py; leave the cmd window open to watch live progress and errors.

File layout (for developers)

enhanced-pdfer/
├── pyproject.toml
├── requirements.txt
├── README.md
├── CHANGES.md
├── MIGRATION.md
└── src/enhanced_pdfer/
    ├── __main__.py        # entrypoint: mode dialog, pickers, server, browser
    ├── app.py             # FastAPI app + stage-gated endpoints
    ├── session.py         # in-memory session state + stage machine + manifest
    ├── pipeline.py        # orchestrator (combine + single + finalize)
    ├── pickers.py         # native OS dialogs (tkinter)
    ├── single_pdf.py      # rotate / extract / split / crop / load (single mode)
    ├── redact.py          # true redaction (rect + regex)
    ├── signature.py       # signature image placement + type-to-sign
    ├── stamps.py          # preset + custom stamps
    ├── ocr.py             # ocrmypdf wrapper (graceful skip)
    ├── tables.py          # pdfplumber + optional camelot -> xlsx
    ├── optimize.py        # compression (standard + aggressive), watermark, password
    ├── merge.py           # merge + bookmarks + metadata + page reorder
    ├── stamp.py           # orientation-aware page numbers
    ├── blank_detect.py    # visual-blank page scanner
    ├── excel_audit.py     # print-area pre-flight
    ├── convert/           # platform conversion dispatch (COM-hardened on Win)
    └── web/               # index.html, app.js, style.css

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages