Turn a stream of scientific literature into validated, audit-ready editorial output — a reproducible pipeline patterned on a production system that processes three journals every month.
This repository demonstrates the sanitized metadata-enrichment layer of that system, with synthetic samples and optional live providers. The full production architecture — six processing stages, cross-language scientific editing, and document generation — is described below as proof of the engineering behind it.
A scientific publisher prepared a recurring digest of international research — a monthly task that took a skilled editor 6–10 hours per journal: locate the relevant articles, translate each abstract, edit it to a scientific standard, apply literary editing, and format the result into the journal's house style.
It was expensive, slow, dependent on scarce expertise, and impossible to scale beyond a single journal without adding people. The kind of recurring, judgment-heavy expert work that quietly consumes the most valuable hours in a specialist organization.
The production system this repository is patterned on replaced that manual task with an automated, auditable pipeline:
| Before | After | |
|---|---|---|
| Time per issue | 6–10 hours | under 1 hour |
| Manual translation / editing passes | 3–5 | 0 (expert kept on final review) |
| Journals covered | 1 (with effort) | 3 simultaneously |
| Quality | depends on operator and day | reproducible (same input → same output) |
| Traceability | none | full audit trail per run |
These outcomes are from the private production system this public repository is patterned on, in production in 2026 with outputs delivered to a real editorial team. The public repo demonstrates the sanitized enrichment layer with synthetic samples — it does not reproduce the production run.
The full system processes each journal issue through six sequential stages, turning a request into finished editorial documents:
[Request] → Source → Filter → Normalize → Translate & Edit → Assemble → [Documents]
1 · Source — retrieves metadata for every article in an issue via the Crossref API. For publishers that do not expose full abstracts there, abstracts are recovered from secondary scholarly sources (Semantic Scholar, Europe PMC).
2 · Filter — removes everything that does not belong in a digest: corrections, retractions, notices, service material, and records without usable abstracts.
3 · Normalize — brings raw data to a single shape: strips HTML, normalizes whitespace, standardizes author lists, and extracts the structural sections of each abstract (Introduction, Methods, Results, Conclusion).
4 · Translate & Edit (the core stage) — each abstract is processed by a language model in a single call governed by an editorial prompt contract that combines scientific translation, editing, and correction. Output passes a deterministic validation and enforcement layer before it is accepted.
5 · Assemble — processed material is rendered into two aligned Word documents per issue: a target-language editorial draft in the journal's house format, and the English source for editorial control, plus a markdown archival copy.
6 · Integrate & Validate — a CLI orchestrator runs all stages as one command. Every run is recorded to a structured log (status, output paths, statistics), and coverage is checked automatically against a golden reference set when available.
The properties that turn an AI experiment into a system an editorial team can rely on:
- 429 automated tests (pytest) covering every stage — all passing.
- Golden reference sets per journal that verify each run covers the full issue, so nothing is silently dropped.
- Translation Quality (TQ) protocol — a structured four-document review cycle grading every translation against critical / major / cosmetic error classes.
- Model comparison tooling — quality of different LLM versions is compared run-to-run before any model change ships.
- Full audit trail — every run is reproducible and logged end to end.
Python 3.13 · Crossref / Semantic Scholar / Europe PMC (metadata) · OpenAI (LLM) · python-docx (document generation) · PyYAML (configuration) · requests · lxml · pytest.
The production system runs on private editorial templates, domain configuration, and client data, and is not included in this public repository. The section below is the sanitized, runnable slice you can inspect and execute yourself.
The public repo is the metadata-enrichment layer of the architecture above, packaged as a runnable, CI-safe demonstration on synthetic data:
DOI list
│
▼
[Fetch] retrieve bibliographic metadata for each DOI
│
▼
[Enrich] add structured AI summary, keywords, and topic tags
│
▼
[Validate] enforce schema on every enriched record
│
▼
[Output] per-record JSON + manifest with SHA-256 checksums
One command turns a list of identifiers into validated, checksummed records ready for downstream ingest — no manual extraction or formatting in between. It runs out of the box with no API keys (the default mock provider is deterministic and CI-safe), so you can verify the architecture and verification pattern in seconds.
This is not a thin wrapper around an LLM. The architecture is designed so the output can be relied on:
- Reproducible — deterministic processing; the same input produces the same output.
- Auditable — every run produces a manifest with per-record SHA-256 checksums.
- Validated — schema enforcement on each record; malformed output fails fast.
- Provider-agnostic — metadata and LLM backends swap behind a stable interface, without touching pipeline logic.
These are the same properties — applied here to the enrichment layer — that make the production system above trustworthy.
git clone https://github.com/DmitryIri/ai-editorial-digest-pipeline.git
cd ai-editorial-digest-pipeline
pip install -e ".[dev]"
python examples/quickstart.pyExpected output:
Enriched 2 records:
DOI: 10.9999/synthetic.2024.001
Title: Automated Knowledge Extraction from Scientific Literature
AI Summary: This paper titled 'Automated Knowledge Extraction from
Scientific Literature' presents research in the domain of
NLP, knowledge extraction.
AI Topics: methodology, empirical study
Provider: mock
Manifest root checksum: 68f60e9684391180...
Records: 2
Runs out of the box with no API keys — the default mock provider is CI-safe and deterministic.
Each enriched record:
{
"doi": "10.9999/synthetic.2024.001",
"title": "Automated Knowledge Extraction from Scientific Literature",
"abstract": "We present a synthetic study on automated extraction of structured knowledge...",
"authors": ["Alice Researcher", "Bob Scholar"],
"year": 2024,
"journal": "Journal of Synthetic Research",
"keywords": ["NLP", "knowledge extraction"],
"ai_summary": "This paper titled 'Automated Knowledge Extraction from Scientific Literature' presents research in the domain of NLP, knowledge extraction.",
"ai_keywords": ["ai-generated", "mock", "NLP", "knowledge extraction", "scientific literature"],
"ai_topics": ["methodology", "empirical study"],
"enrichment_provider": "mock"
}Manifest with reproducibility checksums:
{
"generated_at": "2024-01-15T10:30:00+00:00",
"record_count": 2,
"items": [{"doi": "10.9999/synthetic.2024.001", "checksum": "sha256..."}],
"root_checksum": "sha256..."
}DOI list
│
▼
[Fetcher] ← METADATA_PROVIDER=mock | crossref
│ RawMetadata (Pydantic)
▼
[Enricher] ← LLM_PROVIDER=mock | openai | anthropic
│ EnrichedMetadata (Pydantic)
▼
[Validator] ← schema enforcement
│
▼
JSON records + Manifest (SHA-256 checksums)
A clean separation of concerns: each stage has a typed contract, so backends can change without rewriting the pipeline.
The same pattern — take an external content stream, turn it into validated, traceable, standardized output — applies wherever expert time is spent on repeatable extraction and structuring:
- Research portals — auto-enrich imported bibliographic records
- Publishing automation — metadata completion and quality checks before ingest
- Literature review / RAG preparation — structured extraction for downstream pipelines
- Knowledge graph population — structured entities from abstracts at scale
pytest tests/ -v
ruff check src/ tests/ examples/ tools/All tests run on the mock provider — no API keys required. CI runs on Python 3.10 / 3.11 / 3.12.
This public repository is the enrichment layer, packaged as a runnable, CI-safe demonstration:
- Default mode is mock — synthetic samples, no API keys, deterministic output. Safe for CI and for running in seconds.
- Live providers are optional integration points —
METADATA_PROVIDER=crossrefandLLM_PROVIDER=openai|anthropicare defined interfaces ready to be wired to real backends. - Sample data in
fixtures/is synthetic — no real publications or authors.
The architecture, contracts, validation, and reproducibility pattern are real; the live API backends are the layer left open for adaptation to a specific client's sources and standards. The translation/editing and document-generation stages of the production system are domain- and template-specific and are not part of this public demonstration.
- Python 3.10+
- Pydantic v2 — data validation and schema enforcement
- httpx — HTTP client for metadata integration
- ruff — linting · pytest + pytest-cov — testing
- GitHub Actions CI (Python 3.10 / 3.11 / 3.12)
MIT