You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
**Give AI agents the power to process, sign, and transform documents.**
11
+
**Give AI agents the power to generate, read, extract, process, and sign documents.**
12
12
13
13
## Description
14
14
15
-
A Model Context Protocol (MCP) server that connects AI assistants to the [Nutrient Document Web Service (DWS) Processor API](https://www.nutrient.io/api) — enabling document creation, editing, conversion, digital signing, OCR, redaction, and more through natural language.
15
+
A Model Context Protocol (MCP) server that connects AI assistants to the [Nutrient Document Web Service (DWS)](https://www.nutrient.io/api)Processor and Data Extraction APIs — enabling document creation, editing, conversion, digital signing, OCR, and redaction, plus structured data extraction (typed JSON with bounding boxes and confidence, or schema-guided field extraction with per-field citations) through natural language.
16
16
17
17
## Features
18
18
19
19
- Local stdio MCP server for Claude Desktop and other MCP-compatible clients
20
20
- Browser-based OAuth on the first request that uses the Nutrient API, with optional API-key fallback for CI and headless environments
21
-
- Document conversion, OCR, extraction, redaction, watermarking, annotation flattening, and digital signing
21
+
- Document conversion, OCR, redaction, watermarking, annotation flattening, and digital signing (Processor API)
22
+
- Data extraction (Data Extraction API): parse whole documents to Markdown or spatial JSON, then pull named fields into a JSON schema you define, with per-field citations. Four parse modes: `text` (1 credit/page, no OCR), `structure` (1.5), `understand` (9, the default), `agentic` (18, VLM)
22
23
- Sandbox-aware local file handling with explicit output paths
23
24
- Read-only account lookup for DWS credits and usage
24
25
@@ -41,6 +42,9 @@ Once configured, you (or your AI agent) can process documents through natural la
41
42
**You:**_"OCR this scanned document in German and extract the text"_
42
43
**AI:**_"I've processed the scan with German OCR. Here's the extracted text..."_
43
44
45
+
**You:**_"Pull the vendor, invoice number, total, and due date out of invoice-0341.pdf, with citations"_
46
+
**AI:**_"Here are the four fields as JSON. Each value cites the page and bounding box it came from..."_
47
+
44
48
## Installation
45
49
46
50
Install it from Claude Desktop Settings -> Extensions if you are using Claude Desktop. If you are developing locally, use the manual setup below.
@@ -266,6 +270,12 @@ These examples assume your files live inside the configured sandbox and that you
266
270
267
271
**What happens:** The server first performs a read-only account lookup, then converts the DOCX file to PDF, saves the result in the sandbox, and tells the user exactly where the output file was written.
268
272
273
+
### Example 4: Schema-guided field extraction with citations
274
+
275
+
**User prompt:**`Extract vendor_name, invoice_number, total_amount and due_date from /path/to/sandbox/invoice-0341.pdf and save the citations next to it.`
276
+
277
+
**What happens:** The agent calls `extract_fields` with a small JSON schema (`{ "type": "object", "properties": { "vendor_name": {"type": "string"}, "invoice_number": {"type": "string"}, "total_amount": {"type": "string"}, "due_date": {"type": "string"} } }`) and an `outputPath`. The server sends the PDF to the Data Extraction API, returns the four values inline as JSON with a citation match summary, and writes the full per-field citations (page, bounding box, confidence) to the output file for auditing.
278
+
269
279
## Use with AI Agent Frameworks
270
280
271
281
This MCP server works with any platform that supports the Model Context Protocol:
0 commit comments