fix(sdk): route XLSX through column-aware pseudonymize_xlsx_smart - #5
Merged
Merged
Conversation
`Redactor.redact_file` (CLI/SDK/daemon) flattened every cell across
every sheet into a single ` | `-joined string before detection, then
did substring `cell.value.replace()` on reconstruction. Cell context
was destroyed and many entities — short surnames, locations,
numerically-typed cells — were silently missed.
The SDK now branches on the XLSX MIME type and calls
`xlsx_inference.pseudonymize_xlsx_smart`, the same column-aware path
the noirdoc-cloud proxy has always used: header-keyword classification,
per-column NLP sampling, and per-cell `<<TYPE_N>>` pseudonyms via
`mapper.get_or_create()`. The reveal path was already cell-aware.
The flat-text `extract_xlsx` helper is retained for
`pipeline.convert_unsupported_files` ("ship XLSX as text to a
non-Excel-aware LLM"); its docstring now documents that redaction
must use `pseudonymize_xlsx_smart` instead.
Adds `tests/test_sdk_xlsx.py` covering redact-on-classified-columns
and the redact→reveal round-trip.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Redactor.redact_file(CLI / SDK / daemon) was flattening every cell across every sheet into a single|-joined string before detection, then doing substringcell.value.replace()on reconstruction. Cell context was destroyed and many entities — short surnames, locations, numerically-typed cells — were silently missed.noirdoc.file_analysis.xlsx_inference.pseudonymize_xlsx_smart— the same column-aware path the noirdoc-cloud proxy has always used (header-keyword classification + per-column NLP sampling + per-cell<<TYPE_N>>pseudonyms viamapper.get_or_create()). The reveal path was already cell-aware and round-trips correctly.extract_xlsxhelper is retained forpipeline.convert_unsupported_files("ship XLSX as text to a non-Excel-aware LLM"); its docstring now documents that redaction must usepseudonymize_xlsx_smartinstead.[0.1.2]section with a### Fixedentry.Test plan
pytest tests/test_sdk_xlsx.py -v— both new tests pass (redact on classified columns; redact→reveal round-trip).pytest tests/ -m "not slow"— 238 passed; the 4 pre-existing PDF errors are test-ordering pollution unrelated to this change (reproduces onmainwithout these edits).noirdoc redact sample.xlsx -o out.xlsxon a German-style workbook withName / E-Mail / Telefon / IBAN / Notizencolumns; confirm classified columns get<<TYPE_N>>tokens andNotizenis untouched.noirdoc reveal out.xlsx -o restored.xlsxand confirm originals are back.