Skip to content

fix(sdk): route XLSX through column-aware pseudonymize_xlsx_smart - #5

Merged
comppaz merged 1 commit into
mainfrom
fix/xlsx-smart-pipeline-from-sdk
Apr 27, 2026
Merged

comppaz merged 1 commit into
mainfrom
fix/xlsx-smart-pipeline-from-sdk

Conversation

@comppaz

@comppaz comppaz commented Apr 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Redactor.redact_file (CLI / SDK / daemon) was flattening every cell across every sheet into a single |-joined string before detection, then doing substring cell.value.replace() on reconstruction. Cell context was destroyed and many entities — short surnames, locations, numerically-typed cells — were silently missed.
  • XLSX inputs now branch on MIME and route through noirdoc.file_analysis.xlsx_inference.pseudonymize_xlsx_smart — the same column-aware path the noirdoc-cloud proxy has always used (header-keyword classification + per-column NLP sampling + per-cell <<TYPE_N>> pseudonyms via mapper.get_or_create()). The reveal path was already cell-aware and round-trips correctly.
  • The flat-text extract_xlsx helper is retained for pipeline.convert_unsupported_files ("ship XLSX as text to a non-Excel-aware LLM"); its docstring now documents that redaction must use pseudonymize_xlsx_smart instead.
  • CHANGELOG: extended the unreleased [0.1.2] section with a ### Fixed entry.

Test plan

  • pytest tests/test_sdk_xlsx.py -v — both new tests pass (redact on classified columns; redact→reveal round-trip).
  • pytest tests/ -m "not slow" — 238 passed; the 4 pre-existing PDF errors are test-ordering pollution unrelated to this change (reproduces on main without these edits).
  • Manual smoke: noirdoc redact sample.xlsx -o out.xlsx on a German-style workbook with Name / E-Mail / Telefon / IBAN / Notizen columns; confirm classified columns get <<TYPE_N>> tokens and Notizen is untouched.
  • noirdoc reveal out.xlsx -o restored.xlsx and confirm originals are back.
  • Daemon parity: repeat the smoke test with the daemon running.

`Redactor.redact_file` (CLI/SDK/daemon) flattened every cell across
every sheet into a single ` | `-joined string before detection, then
did substring `cell.value.replace()` on reconstruction. Cell context
was destroyed and many entities — short surnames, locations,
numerically-typed cells — were silently missed.

The SDK now branches on the XLSX MIME type and calls
`xlsx_inference.pseudonymize_xlsx_smart`, the same column-aware path
the noirdoc-cloud proxy has always used: header-keyword classification,
per-column NLP sampling, and per-cell `<<TYPE_N>>` pseudonyms via
`mapper.get_or_create()`. The reveal path was already cell-aware.

The flat-text `extract_xlsx` helper is retained for
`pipeline.convert_unsupported_files` ("ship XLSX as text to a
non-Excel-aware LLM"); its docstring now documents that redaction
must use `pseudonymize_xlsx_smart` instead.

Adds `tests/test_sdk_xlsx.py` covering redact-on-classified-columns
and the redact→reveal round-trip.
@comppaz
comppaz merged commit 69ff4ce into main Apr 27, 2026
3 checks passed
@comppaz
comppaz deleted the fix/xlsx-smart-pipeline-from-sdk branch April 27, 2026 14:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant