Skip to content

Commit bbf749e

Browse files
authored
docs: surface Data Extraction API in README copy and header image (#40)
1 parent bddb64f commit bbf749e

2 files changed

Lines changed: 14 additions & 4 deletions

File tree

‎README.md‎

Lines changed: 14 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -6,19 +6,20 @@
66
<img width="380" height="200" src="https://glama.ai/mcp/servers/@PSPDFKit/nutrient-dws-mcp-server/badge" alt="Nutrient DWS MCP Server" />
77
</a>
88

9-
[![npm](https://img.shields.io/npm/v/%40nutrient-sdk/dws-mcp-server)](https://www.npmjs.com/package/@nutrient-sdk/dws-mcp-server)
9+
[![npm](https://img.shields.io/npm/v/%40nutrient-sdk/dws-mcp-server)](https://www.npmjs.com/package/@nutrient-sdk/dws-mcp-server) [![smithery badge](https://smithery.ai/badge/nutrient/dws-mcp-server)](https://smithery.ai/servers/nutrient/dws-mcp-server)
1010

11-
**Give AI agents the power to process, sign, and transform documents.**
11+
**Give AI agents the power to generate, read, extract, process, and sign documents.**
1212

1313
## Description
1414

15-
A Model Context Protocol (MCP) server that connects AI assistants to the [Nutrient Document Web Service (DWS) Processor API](https://www.nutrient.io/api) — enabling document creation, editing, conversion, digital signing, OCR, redaction, and more through natural language.
15+
A Model Context Protocol (MCP) server that connects AI assistants to the [Nutrient Document Web Service (DWS)](https://www.nutrient.io/api) Processor and Data Extraction APIs — enabling document creation, editing, conversion, digital signing, OCR, and redaction, plus structured data extraction (typed JSON with bounding boxes and confidence, or schema-guided field extraction with per-field citations) through natural language.
1616

1717
## Features
1818

1919
- Local stdio MCP server for Claude Desktop and other MCP-compatible clients
2020
- Browser-based OAuth on the first request that uses the Nutrient API, with optional API-key fallback for CI and headless environments
21-
- Document conversion, OCR, extraction, redaction, watermarking, annotation flattening, and digital signing
21+
- Document conversion, OCR, redaction, watermarking, annotation flattening, and digital signing (Processor API)
22+
- Data extraction (Data Extraction API): parse whole documents to Markdown or spatial JSON, then pull named fields into a JSON schema you define, with per-field citations. Four parse modes: `text` (1 credit/page, no OCR), `structure` (1.5), `understand` (9, the default), `agentic` (18, VLM)
2223
- Sandbox-aware local file handling with explicit output paths
2324
- Read-only account lookup for DWS credits and usage
2425

@@ -41,6 +42,9 @@ Once configured, you (or your AI agent) can process documents through natural la
4142
**You:** _"OCR this scanned document in German and extract the text"_
4243
**AI:** _"I've processed the scan with German OCR. Here's the extracted text..."_
4344

45+
**You:** _"Pull the vendor, invoice number, total, and due date out of invoice-0341.pdf, with citations"_
46+
**AI:** _"Here are the four fields as JSON. Each value cites the page and bounding box it came from..."_
47+
4448
## Installation
4549

4650
Install it from Claude Desktop Settings -> Extensions if you are using Claude Desktop. If you are developing locally, use the manual setup below.
@@ -266,6 +270,12 @@ These examples assume your files live inside the configured sandbox and that you
266270

267271
**What happens:** The server first performs a read-only account lookup, then converts the DOCX file to PDF, saves the result in the sandbox, and tells the user exactly where the output file was written.
268272

273+
### Example 4: Schema-guided field extraction with citations
274+
275+
**User prompt:** `Extract vendor_name, invoice_number, total_amount and due_date from /path/to/sandbox/invoice-0341.pdf and save the citations next to it.`
276+
277+
**What happens:** The agent calls `extract_fields` with a small JSON schema (`{ "type": "object", "properties": { "vendor_name": {"type": "string"}, "invoice_number": {"type": "string"}, "total_amount": {"type": "string"}, "due_date": {"type": "string"} } }`) and an `outputPath`. The server sends the PDF to the Data Extraction API, returns the four values inline as JSON with a citation match summary, and writes the full per-field citations (page, bounding box, confidence) to the output file for auditing.
278+
269279
## Use with AI Agent Frameworks
270280

271281
This MCP server works with any platform that supports the Model Context Protocol:

‎resources/readme-header.png‎

60.2 KB
Loading

0 commit comments

Comments
 (0)