This application anonymizes large PDF, Markdown, or Text files using a hybrid approach: a fast, safe RE2-powered regex pre-filter (google-re2) followed by LLM-based semantic detection.
- High-Quality Anonymization: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
- Fast & Safe Regex Pre-Filter (Hybrid NER): First stage uses the RE2 engine (
google-re2) for linear-time, ReDoS-safe detection of 70+ structured PII patterns (emails, phones, URLs, credit cards, IBANs, crypto wallets, VINs, MACs, dates, currency, plus national IDs/tax IDs/driver licenses/VAT/business/government IDs etc.). Country-partitioned patterns cover 30+ jurisdictions (mandatory: USA, Canada, UK, Spain, Italy, France, India, China + DE, JP, BR, AU, NL, and many more). Regex results are priority-merged with the LLM stage. Full list and per-country details are in the core package docs and recipes. - Large File Support: Consistently anonymizes large files (tested up to 1GB).
- Multi-Provider & Cost-Effective: Free to use with local Ollama models. It also supports major providers like OpenAI, Anthropic, Google, Hugging Face, and OpenRouter.
- Reversible: Supports deanonymization to recover original data when needed.
- Multi-Format: Works with PDF, Markdown, and plain text files.
A comprehensive documentation site is available at leo-gan.github.io/anonymizer/.
The documentation includes:
- Anonymization 101: A guide on data anonymization and deanonymization techniques.
- Installation Guide: System requirements, package extras, and setup.
- CLI Usage: Reference manual for
pdf-anonymizer runandpdf-anonymizer deanonymize(including configuration profiles). - API Usage: Programmatic usage guide for the core SDK.
- API Reference (auto): Auto-generated function signatures and details.
- Recipes & Common Workflows: Practical patterns (local Ollama, external LLM round-trips, batching, profiles, caching, debugging, etc.).
- Troubleshooting: Common issues and solutions.
- Architecture: Understanding how the anonymizer operates.
This project is a monorepo containing two main packages:
packages/pdf-anonymizer-core: The core library containing the anonymization and deanonymization logic. See the core README for more details.packages/pdf-anonymizer-cli: A command-line interface for using the anonymizer. See the CLI README for detailed usage instructions.
-
Install
uv: This project usesuvfor package management. Follow the official installation instructions. -
Clone the repository:
git clone https://github.com/leo-gan/anonymizer.git cd anonymizer -
Install dependencies:
uv sync --group dev
-
Install Ollama (optional): If you want to use a local model for anonymization, install Ollama.
-
Set up environment variables: Create a
.envfile in the root directory of the repository (or in thepackages/pdf-anonymizer-clidirectory) and add the necessary API keys for the providers you want to use. For example:# For Google models GOOGLE_API_KEY="YOUR_GOOGLE_API_KEY" # For OpenAI models OPENAI_API_KEY="YOUR_OPENAI_API_KEY" # For Anthropic models ANTHROPIC_API_KEY="YOUR_ANTHROPIC_API_KEY" # For Hugging Face models HUGGING_FACE_TOKEN="YOUR_HF_TOKEN" # For OpenRouter models OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
To anonymize a file, run the pdf-anonymizer CLI tool using uv run:
uv run pdf-anonymizer run document.pdfCommonly you pick a configuration profile with -p / --config-profile (or let it default to best-speed):
uv run pdf-anonymizer run contract.pdf -p best-quality
uv run pdf-anonymizer run large-notes.md -p best-cost --model-name "ollama/phi4-mini"To deanonymize the file later:
uv run pdf-anonymizer deanonymize document.anonymized.md document.mapping.jsonFor detailed command-line options and examples (including all three profiles and overrides), please refer to the CLI README, the CLI Usage Docs, or the Recipes & Common Workflows page.
To demonstrate the hybrid NER (RE2 regex pre-filter + LLM) and the complete round-trip process on a real document, we have provided a demo script:
-
Prepare the Demo Document: This script downloads an open-access arXiv research paper PDF and injects synthetic PII (name, email, phone, IP, SSN) onto the first page:
uv run python scripts/prepare_demo_pdf.py
-
Run the Demo: This script runs the hybrid NER (fast RE2-based regex for structured PII + LLM semantic stage) on the PDF, anonymizes the PII to structured placeholders, saves the mapping vocabulary, and performs a complete round-trip deanonymization:
uv run python scripts/demo_anonymize.py
You will see colorized console logs showing the exact matched entities (including the rich set of regex-detected PII types), the anonymized text, and the fully-reverted deanonymized output.
The built-in regex stage now automatically detects a wide range of additional structured PII (credit cards, IBANs, crypto addresses, VINs, MAC addresses, many country-specific national/tax/driver IDs, etc.) on top of classic items like email/phone/IP/SSN.
To run the test suite:
uv run pytest- Documentation Site — Full guides, 101 course, and API reference.
- Core Package README — SDK details and installation.
- CLI Package README — Command-line specifics.
- Recipes & Common Workflows — Practical usage examples.
- Troubleshooting — Solutions to common problems.