Skip to content
#

scanned-documents

Here are 72 public repositories matching this topic...

Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser

  • Updated Sep 15, 2026
  • Python

BoxDetect is a Python package based on OpenCV which allows you to easily detect rectangular shapes like character or checkbox boxes on scanned forms.

  • Updated Jan 18, 2023
  • Python

AI-agent skill producing reusable Markdown from PDFs. It turns flowcharts, diagrams, and charts into text beside each caption instead of empty links. It checks an earlier conversion against the PDF and fixes misread or missing parts. Long PDFs run in small saved batches with an independent review pass, each paragraph tagged with its page.

  • Updated Aug 23, 2026
  • Python

Add this topic to your repo

To associate your repository with the scanned-documents topic, visit your repo's landing page and select "manage topics."

Learn more