Smart local file search app that understands your files
-
Updated
Apr 10, 2026 - Python
Smart local file search app that understands your files
可以将word(doc、docx)、excel、pdf、ppt、csv、txt文件的文本内容提取出来,同时能够提取出word、pdf文件的目录
A suite of Machine Learning / Deep Learning Dockerfiles to allow Apache Tika to extract objects and to produce textual captions for images and video
tokyo, a REST API, when given any type of document 📄, Identifies mime-type 🧐. Suggests extension 🦔. Alas Extracts text 💪.
Extract text from a document by Apache Tika
AWS Lambda layer containing latest version of Apache Tika
Text extraction from scanned pdf documents in java
The metadata and text content extractor for almost every file type.
Visualize unstructured data using Watson NLU
A permissively licensed crate to detect MIME types
ApacheDeepLearning101
Apache NiFi + Apache Tika + OptimaizeLangDetector
All my processors (NARs) in one place
🚴♂️⛷Data Lake, Performance tuning for text extraction from a huge amount of files.
CLI keyword search across websites, documents, and local folders. Crawls multi-level sites, renders JavaScript SPAs with a headless browser, and extracts text from PDF, DOCX and 100+ formats via Apache Tika — reporting exact page and line numbers. Fuzzy matching, JSON output, Docker image.
Directory tree metadata parser using Apache Tika
Custom search engine for all kinds of documents and storage services
Developed a Spatial Search website that allow users to search documents from FBI Vault website. Extract the most frequently occurring location in each of documents, and load the geo-tagged data into Apache Solr to index the documents, visualize search results using the Google Maps API.
To associate your repository with the apache-tika topic, visit your repo's landing page and select "manage topics."