Simplifies the retrieval, extraction, and training of structured data from various unstructured sources.
-
Updated
Oct 22, 2025 - Python
Simplifies the retrieval, extraction, and training of structured data from various unstructured sources.
Structured data extraction from research literature
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
A ruby gem to extract structured data from Google Local Search Results using the serpapi/bert-base-local-results model, enabling parsing, classification, and information extraction from English HTML content.
Extract structured data from semi-structured documents using any OpenAI-compatible LLM, with per-field grounding and confidence.
Convert any document format into LLM-ready data format (markdown) with advanced intelligent document processing capabilities powered by pre-trained models.
find a template of many similar html files
A new package facilitates extracting a concise, structured summary from user-provided news headlines or brief texts by utilizing pattern matching and LLM interactions. This tool aims to help researche
A new package that leverages language models to transform structured YAML data into well-formatted resume PDFs. Users provide their resume details in YAML format, and the package extracts key informat
This new package facilitates extracting structured insights from text-based content related to domain-specific issues, such as analyzing DNS blocking reports. Given unstructured text describing networ
A new package is designed to analyze financial news headlines and extract key structured information such as company names, financial targets, timeframes, and goal updates from text inputs. It simplif
A new package designed to facilitate structured extraction of key information from scientific or factual text inputs, enabling precise summaries, data extraction, or categorization based on user promp
A working AI agent that triages invoices end to end on LangGraph: extracts fields from raw text, validates against business rules, self-corrects, and posts to a mock ERP. Includes a held-out eval with real numbers. A winder.ai demo.
A new package that analyzes technical arguments and extracts structured summaries from text discussions about infrastructure-as-code practices. It takes user-provided text (such as forum posts, articl
A complete engineering case study in building an event-driven OCR pipeline: system design, distributed workers, hybrid AI model routing, and production observability - not just an OCR API wrapper.
SchemaSure — independent third-party profile of a public API surface, by API Evangelist. Structured data extraction API that converts unstructured text or HTML into JSON guaranteed to validate against a caller-supplied JSON Schema, or returns a typed error at no charge. Access is gated by x402 pay-per-call micropayments (USDC on Base), with no acco
Performance driven Go library for crawling and extracting structured content from the web.
LlamaParse — independent third-party profile of a public API surface, by API Evangelist. LlamaParse is an enterprise document parsing and AI pipeline platform from LlamaIndex that converts complex PDFs, Office files, and 130+ document formats into LLM-ready structured outputs.
A Python-based tool for extracting structured data from PDFs using OCR and regex, and exporting it to CSV. Ideal for processing invoices, logs, or scanned documents into organized, usable datasets.
Supports text, image and structured data
To associate your repository with the structured-data-extraction topic, visit your repo's landing page and select "manage topics."