Skip to content

Repository files navigation

MinerUNativeExtract

Lean standalone Swift CLI for notemd PDF extraction.

The package extracts:

  • embedded PDF text into document.md
  • page JPEGs
  • cropped JPEG assets for images, charts, and tables
  • manifest.json with page/block metadata

Formula OCR/LaTeX extraction is intentionally not included. Formula layout classes from detector models are ignored.

Build

swift build -c release

The binary is written to:

.build/release/mineru-native-extract

Run

Recommended standalone layout path:

swift run mineru-native-extract \
  --path input.pdf \
  --output extraction-dir \
  --layout-onnx-model Models/PP-DocLayout-M.onnx \
  --layout-labels Models/PP-DocLayout-M-labels.json \
  --layout-input-size 640x640 \
  --layout-confidence 0.3 \
  --require-layout-source

Debug layout overlays:

swift run mineru-native-extract \
  --path input.pdf \
  --output extraction-dir \
  --layout-onnx-model Models/PP-DocLayout-M.onnx \
  --layout-labels Models/PP-DocLayout-M-labels.json \
  --layout-debug \
  --require-layout-source

Optional Core ML delegation for ONNX Runtime:

swift run mineru-native-extract \
  --path input.pdf \
  --output extraction-dir \
  --layout-onnx-model Models/PP-DocLayout-M.onnx \
  --layout-labels Models/PP-DocLayout-M-labels.json \
  --onnx-coreml

--onnx-coreml may or may not improve speed depending on the model and machine. The CPU path remains the default.

Output

extraction-dir/
  document.md
  manifest.json
  assets/
    page-001.jpg
    page-001-image-003.jpg
    page-002-chart-004.jpg
    page-003-table-002.jpg
  debug/
    page-001-layout.jpg

debug/ exists only when --layout-debug is enabled.

Runtime Contract

This package does not require MinerU or Python at runtime. The preferred runtime bundle for notemd is:

mineru-native-extract
Models/PP-DocLayout-M.onnx
Models/PP-DocLayout-M-labels.json

The PP-DocLayout-M conversion/download helpers are in:

Tools/PPDocLayoutM/

CLI Reference

--path input.pdf
--output extraction-dir
--dpi 200
--jpeg-quality 0.88

--layout-onnx-model PP-DocLayout-M.onnx
--layout-labels labels.json
--layout-input-name image
--layout-scale-factor-input-name scale_factor
--layout-detections-output-name fetch_name_0
--layout-input-size 640x640
--layout-confidence 0.3

--layout-coreml-model LayoutDetector.mlmodelc
--layout-boxes-output-name boxes
--layout-scores-output-name scores
--layout-classes-output-name classes
--layout-detections-format class-score-xyxy

--mineru-middle-json file.json
--require-layout-source
--layout-debug
--onnx-coreml

Use only one layout source at a time: ONNX, Core ML, or MinerU middle JSON.

About

PDF extraction for science in Swift

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages