Agent skill for pixel-grounded chart data extraction
-
Updated
Jul 3, 2026 - Python
Agent skill for pixel-grounded chart data extraction
Python library for extracting content from PowerPoint files including embedded charts and SmartArt. Built for RAG and document processing pipelines.
CUDA-accelerated PDF -> Markdown/HTML converter using Docling + IBM Granite Vision chart extraction
Extract charts, figures, and tables from academic PDFs for AI agent analysis
A complete end-to-end pipeline for extracting structured data from chart and graph images
Converters where figures survive. DOCX/XLSX to Markdown via native OOXML chart data: real numbers, OCR/VLM are only optional. CLI, MCP server, Docker image, and .mcpb bundle included.
doc-textify: offline, CPU-only PDF/image to Markdown & LLM-ready text converter. OCR with CJK normalization, layout recovery, two-column reading order, table/chart/formula extraction, RAG-ready chunking — no vision LLMs, no GPU, no cloud.
Add a description, image, and links to the chart-extraction topic page so that developers can more easily learn about it.
To associate your repository with the chart-extraction topic, visit your repo's landing page and select "manage topics."