The model only has one job here: turn a document into the JSON schema in
asienta/extraction/schema.py. Everything else is rules, so any model that can read an image
or a PDF and follow a schema will do. Pick one in config.ini; nothing else changes.
| Provider | [reader] |
Key | Notes |
|---|---|---|---|
| Google Gemini | provider = geminimodel = gemini-3.8-flash |
GEMINI_API_KEY |
Plain REST, no extra package. Flash models are the cheapest cloud option for invoices. |
| Anthropic Claude | provider = claudemodel = claude-opus-5 (or claude-sonnet-5, claude-haiku-4-5) |
ANTHROPIC_API_KEY |
pip install "asienta[claude]". Reads PDFs natively, structured outputs guarantee the schema. |
| OpenAI | provider = openaimodel = <a vision model> |
OPENAI_API_KEY |
Chat Completions with json_schema; PDFs are sent as files. |
| Any OpenAI-compatible API — Azure OpenAI, Mistral, OpenRouter, Together, Groq… | provider = openaibase_url = <their /v1 URL>model = … |
OPENAI_API_KEY (their key) |
Same code, different URL. |
| Local models — Ollama, LM Studio, vLLM | provider = ollamamodel = llama3.2-vision (any vision model you pulled)base_url = http://localhost:11434/v1 |
none | Invoices never leave your network. PDFs are rendered to images first (pip install pymupdf). |
| Demo | provider = demo |
none | Replays stored readings of the sample invoices. |
Accuracy on your documents is what matters, and the cost of a wrong value depends on whether
anything warns about it. asienta bench reads a folder of your invoices with one or more models,
runs the same proposal and checks as the app, and compares the result with a truth file you wrote
by hand:
asienta bench my_invoices/ --truth truth.json --models gemini-3.8-flash,gemini-3.5-flash-liteIt reports, per model: fields read right, perfect invoices, supplier and accounts proposed right,
cost and time per document, and silent errors — wrong values with no warning. A model with a
few visible errors and zero silent ones is safer than a slightly more accurate one that fails
quietly. The truth format is documented at the top of asienta/bench.py; the demo has one
(asienta bench asienta/demo/invoices --truth asienta/demo/truth.json --demo).
Results on 27 real supplier invoices (162 fields), checked by hand, September 2026:
| Model | Fields right | Invoices perfect | Silent errors | Supplier right | Accounts right | Cost / invoice | Time / invoice |
|---|---|---|---|---|---|---|---|
| gemini-3.8-flash (default) | 160/162 | 25/27 | 1 | 27/27 | 24/26 | 0.9 ¢ | 6 s |
| gemini-3.7-flash | 160/162 | 25/27 | 1 | 27/27 | 25/26 | 0.8 ¢ | 6 s |
| gemini-3.5-flash | 160/162 | 25/27 | 0 | 27/27 | 25/26 | 3.5 ¢ | 13 s |
| gemini-3.5-flash-lite | 156/162 | 21/27 | 2 | 27/27 | 24/26 | 0.3 ¢ | 3 s |
| gemini-3.1-flash-lite | 160/162 | 25/27 | 0 | 27/27 | 24/26 | 0.2 ¢ | 4 s |
Other providers weren't part of this run. If you bench one, a pull request adding its row is welcome.
Rules of thumb from real use:
- Thermal tickets and phone photos are where models differ most (digits like 8/6, 1/7). The tax-ID check digit catches most of those, whatever the model.
- Small local models read clean PDFs well and struggle more with crumpled photos; benchmark them on your worst documents before switching.
- Use paid tiers for real invoices: some providers may use free-tier data for training.
The app shows the cost of each reading and the month's total. Prices for common models are
built in; set [reader] price_input and price_output ($ per million tokens) for others.
A reader is one class with read(); see extending.