Skip to content

feat(binary): any-binary normalization stage with distroless-friendly Office conversion - #6

Merged
ancongui merged 2 commits into
mainfrom
feat/binary-normalization
May 14, 2026
Merged

ancongui merged 2 commits into
mainfrom
feat/binary-normalization

Conversation

@ancongui

Copy link
Copy Markdown
Contributor

Summary

Phase 0 of the bbox refinement initiative. The IDP service now accepts any binary callers can throw at it — DOCX, XLSX, PPTX, RTF, ODT, HTML, HEIC/HEIF/AVIF, multi-frame TIFF, SVG, ZIP/7z/TAR bundles, EML/MSG email with attachments — and turns each one into LLM-renderable bytes before extraction. Multilingual by construction: Office docs preserve their original script through Gotenberg/LibreOffice; archives + email pass content through unchanged; the multimodal LLM handles the rest.

Today the loader silently passes unknown formats through to the provider, which fails. This change makes the contract real.

Architecture

  • core/services/binary/ — new package, one adapter per format class:
    • PdfGuard — encrypted/corrupt PDF detection
    • ImageNormalizer — Pillow + pillow-heif (HEIC/HEIF/AVIF), multi-frame TIFF → multi-page PDF, SVG via cairosvg, BMP → PNG
    • OfficeConverter protocol with two adapters:
      • GotenbergConverter (HTTP) — default, distroless-friendly: API + worker containers carry no soffice, just POST bytes to a Gotenberg sidecar
      • LibreOfficeConverter (subprocess) — fallback for slim/dev images that bundle LibreOffice
    • ArchiveUnpacker — ZIP + 7z + TAR + GZIP with zip-bomb guards
    • EmailUnpacker — EML (stdlib) + MSG (extract-msg) → attachments + inline body
    • BinaryNormalizer — orchestrates them all, recurses into archives/emails up to a configurable depth, fans out into the existing documents[] shape
  • Pluggable office adapter wired through pyfly DI: pick via FLYDESK_IDP_OFFICE_CONVERTER=gotenberg|libreoffice. Adding a new adapter (OnlyOfficeConverter, CollaboraConverter, ...) is one new bean + one new enum value.
  • Typed errors (EncryptedPdfError, OfficeConversionError, ArchiveExtractionError, ImageConversionError, UnsupportedBinaryError) all map to RFC 7807 422 problem-details via ExceptionAdvice with stable code strings.
  • Multi-doc bundles (a ZIP, an email with attachments) fan out into multiple per-attachment slots — no API surface change. derived_from carries the unpack chain for traceability.

Plumbing

  • PipelineOrchestrator._step_load calls the normalizer per inbound file before building _FileSlot rows.
  • docker-compose.yml grows a gotenberg service the api + worker depend on. Flip FLYDESK_IDP_OFFICE_CONVERTER=libreoffice to switch back to the in-container subprocess.
  • Dockerfile runtime stage adds libheif1 + libcairo2 + libpango + libgdk-pixbuf so HEIC/SVG conversion works without bloating the image with LibreOffice.
  • New IDPSettings knobs: office_converter, gotenberg_url, gotenberg_timeout_s, binary_normalize_enabled (kill switch), binary_max_recursion_depth, binary_max_expanded_files, binary_libreoffice_path, binary_libreoffice_timeout_s.

Why distroless

Office conversion is the only place subprocess + filesystem writes are forced — soffice won't take stdin/stdout for Office formats and needs a writable user-profile dir. Pure-Python alternatives (python-docx, openpyxl) lose layout fidelity, which would break the Phase 1 bbox refiner downstream. The Gotenberg sidecar is the canonical pattern: bytes-in, PDF-out over HTTP, with soffice + headless Chromium isolated in its own container. The API + worker containers stay distroless-eligible.

Test plan

  • pytest tests/unit → 147 passed, 1 skipped (the SVG test skips on macOS dev machines without libcairo; runs in the Dockerised image where libcairo2 is installed)
  • ruff check . clean
  • pyright src/flydesk_idp 0 errors
  • 53 new unit tests across sniffer, pdf guard, image, archive, email, normalizer (incl. ZIP/EML fan-out + recursion limits), gotenberg (HTTP mocked via respx)
  • Integration smoke against docker compose up with real Gotenberg (covered by Phase 0 follow-up)

Phase plan

This is Phase 0 of the three-phase bbox-refinement initiative:

  1. Phase 0 (this PR) — Binary normalization. Independent of bbox; also fixes today's silent failure on DOCX/EML/ZIP/HEIC.
  2. Phase 1 — Sync grounded bboxes (PyMuPDF + PaddleOCR + script-aware matcher).
  3. Phase 2 — Async out-of-band refinement with PARTIAL_SUCCEEDED job state + new EDA destination.

Multilingual + any-binary are hard requirements across all three phases.

ancongui added 2 commits May 15, 2026 00:36
… Office conversion

Turns every caller-supplied binary into one or more LLM-renderable rows
before the rest of the pipeline runs. PDFs and provider-native rasters
(PNG/JPEG/GIF/WebP) pass straight through; everything else is normalised
on our side so callers can submit DOCX, XLSX, PPTX, RTF, ODT, HTML,
HEIC/HEIF/AVIF iPhone photos, multi-frame TIFF fax scans, SVG,
ZIP/7z/TAR bundles, and EML/MSG email -- including all the multilingual
content their text contains.

Architecture:
* New ``core/services/binary`` package with one adapter per format class
  (PdfGuard, ImageNormalizer, OfficeConverter protocol, ArchiveUnpacker,
  EmailUnpacker), orchestrated by ``BinaryNormalizer``.
* Office conversion is pluggable behind ``OfficeConverter``: the default
  ``GotenbergConverter`` (HTTP) keeps the runtime container distroless-
  friendly by delegating to a Gotenberg sidecar; ``LibreOfficeConverter``
  shells out to local ``soffice`` for slim/dev images that bundle it.
  Selected by ``FLYDESK_IDP_OFFICE_CONVERTER`` (default ``gotenberg``).
* All adapters run through pyfly DI -- ``@service`` autoscan for the
  pure-Python ones, factory ``@bean`` in ``IDPCoreConfiguration`` for
  the OfficeConverter pick + the BinaryNormalizer assembly.
* Typed ``BinaryNormalizationError`` subclasses (encrypted PDF,
  archive extraction failed, office conversion failed, image conversion
  failed, unsupported binary) map to RFC 7807 422s via
  ``ExceptionAdvice``.
* Multi-doc bundles (ZIP, EML+attachments) fan out into the existing
  ``documents[]`` shape -- no API surface change. ``derived_from``
  carries the unpack chain for traceability. Recursion depth + total
  fan-out are bounded by IDPSettings to guard against zip bombs.

Plumbing:
* Orchestrator ``_step_load`` calls the normalizer per inbound file
  before building ``_FileSlot`` rows.
* docker-compose grows a ``gotenberg`` service that the api + worker
  depend on; a single ``FLYDESK_IDP_OFFICE_CONVERTER=libreoffice`` env
  flip switches both back to the in-container subprocess path.
* Dockerfile runtime stage adds libheif1 + libcairo2 + libpango +
  libgdk-pixbuf so HEIC/SVG conversion works without LibreOffice
  bloating the image.

Tests: 53 new unit tests across sniffer, pdf guard, image, archive,
email, normalizer, gotenberg (HTTP mocked via respx). Full suite stays
green; pyright + ruff clean.
@ancongui
ancongui merged commit 04a0a3e into main May 14, 2026
4 checks passed
@ancongui
ancongui deleted the feat/binary-normalization branch May 14, 2026 22:43
ancongui added a commit that referenced this pull request May 31, 2026
… Office conversion (#6)

* feat(binary): any-binary normalization stage with distroless-friendly Office conversion

Turns every caller-supplied binary into one or more LLM-renderable rows
before the rest of the pipeline runs. PDFs and provider-native rasters
(PNG/JPEG/GIF/WebP) pass straight through; everything else is normalised
on our side so callers can submit DOCX, XLSX, PPTX, RTF, ODT, HTML,
HEIC/HEIF/AVIF iPhone photos, multi-frame TIFF fax scans, SVG,
ZIP/7z/TAR bundles, and EML/MSG email -- including all the multilingual
content their text contains.

Architecture:
* New ``core/services/binary`` package with one adapter per format class
  (PdfGuard, ImageNormalizer, OfficeConverter protocol, ArchiveUnpacker,
  EmailUnpacker), orchestrated by ``BinaryNormalizer``.
* Office conversion is pluggable behind ``OfficeConverter``: the default
  ``GotenbergConverter`` (HTTP) keeps the runtime container distroless-
  friendly by delegating to a Gotenberg sidecar; ``LibreOfficeConverter``
  shells out to local ``soffice`` for slim/dev images that bundle it.
  Selected by ``FLYDESK_IDP_OFFICE_CONVERTER`` (default ``gotenberg``).
* All adapters run through pyfly DI -- ``@service`` autoscan for the
  pure-Python ones, factory ``@bean`` in ``IDPCoreConfiguration`` for
  the OfficeConverter pick + the BinaryNormalizer assembly.
* Typed ``BinaryNormalizationError`` subclasses (encrypted PDF,
  archive extraction failed, office conversion failed, image conversion
  failed, unsupported binary) map to RFC 7807 422s via
  ``ExceptionAdvice``.
* Multi-doc bundles (ZIP, EML+attachments) fan out into the existing
  ``documents[]`` shape -- no API surface change. ``derived_from``
  carries the unpack chain for traceability. Recursion depth + total
  fan-out are bounded by IDPSettings to guard against zip bombs.

Plumbing:
* Orchestrator ``_step_load`` calls the normalizer per inbound file
  before building ``_FileSlot`` rows.
* docker-compose grows a ``gotenberg`` service that the api + worker
  depend on; a single ``FLYDESK_IDP_OFFICE_CONVERTER=libreoffice`` env
  flip switches both back to the in-container subprocess path.
* Dockerfile runtime stage adds libheif1 + libcairo2 + libpango +
  libgdk-pixbuf so HEIC/SVG conversion works without LibreOffice
  bloating the image.

Tests: 53 new unit tests across sniffer, pdf guard, image, archive,
email, normalizer, gotenberg (HTTP mocked via respx). Full suite stays
green; pyright + ruff clean.

* style: ruff format pass on Phase 0 binary normalizer files

---------

Co-authored-by: ancongui <andres.contreras@soon.es>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant