Skip to content

Repository files navigation

AI Editorial Digest Pipeline

Turn a stream of scientific literature into validated, audit-ready editorial output — a reproducible pipeline patterned on a production system that processes three journals every month.

This repository demonstrates the sanitized metadata-enrichment layer of that system, with synthetic samples and optional live providers. The full production architecture — six processing stages, cross-language scientific editing, and document generation — is described below as proof of the engineering behind it.

CI Python 3.10+ License: MIT


The Business Problem

A scientific publisher prepared a recurring digest of international research — a monthly task that took a skilled editor 6–10 hours per journal: locate the relevant articles, translate each abstract, edit it to a scientific standard, apply literary editing, and format the result into the journal's house style.

It was expensive, slow, dependent on scarce expertise, and impossible to scale beyond a single journal without adding people. The kind of recurring, judgment-heavy expert work that quietly consumes the most valuable hours in a specialist organization.


From a Manual Bottleneck to a Reproducible Pipeline

The production system this repository is patterned on replaced that manual task with an automated, auditable pipeline:

Before After
Time per issue 6–10 hours under 1 hour
Manual translation / editing passes 3–5 0 (expert kept on final review)
Journals covered 1 (with effort) 3 simultaneously
Quality depends on operator and day reproducible (same input → same output)
Traceability none full audit trail per run

These outcomes are from the private production system this public repository is patterned on, in production in 2026 with outputs delivered to a real editorial team. The public repo demonstrates the sanitized enrichment layer with synthetic samples — it does not reproduce the production run.


The Production System

The full system processes each journal issue through six sequential stages, turning a request into finished editorial documents:

[Request] → Source → Filter → Normalize → Translate & Edit → Assemble → [Documents]

1 · Source — retrieves metadata for every article in an issue via the Crossref API. For publishers that do not expose full abstracts there, abstracts are recovered from secondary scholarly sources (Semantic Scholar, Europe PMC).

2 · Filter — removes everything that does not belong in a digest: corrections, retractions, notices, service material, and records without usable abstracts.

3 · Normalize — brings raw data to a single shape: strips HTML, normalizes whitespace, standardizes author lists, and extracts the structural sections of each abstract (Introduction, Methods, Results, Conclusion).

4 · Translate & Edit (the core stage) — each abstract is processed by a language model in a single call governed by an editorial prompt contract that combines scientific translation, editing, and correction. Output passes a deterministic validation and enforcement layer before it is accepted.

5 · Assemble — processed material is rendered into two aligned Word documents per issue: a target-language editorial draft in the journal's house format, and the English source for editorial control, plus a markdown archival copy.

6 · Integrate & Validate — a CLI orchestrator runs all stages as one command. Every run is recorded to a structured log (status, output paths, statistics), and coverage is checked automatically against a golden reference set when available.

Engineered for trust at production scale

The properties that turn an AI experiment into a system an editorial team can rely on:

  • 429 automated tests (pytest) covering every stage — all passing.
  • Golden reference sets per journal that verify each run covers the full issue, so nothing is silently dropped.
  • Translation Quality (TQ) protocol — a structured four-document review cycle grading every translation against critical / major / cosmetic error classes.
  • Model comparison tooling — quality of different LLM versions is compared run-to-run before any model change ships.
  • Full audit trail — every run is reproducible and logged end to end.

Production stack

Python 3.13 · Crossref / Semantic Scholar / Europe PMC (metadata) · OpenAI (LLM) · python-docx (document generation) · PyYAML (configuration) · requests · lxml · pytest.

The production system runs on private editorial templates, domain configuration, and client data, and is not included in this public repository. The section below is the sanitized, runnable slice you can inspect and execute yourself.


What This Public Repository Demonstrates

The public repo is the metadata-enrichment layer of the architecture above, packaged as a runnable, CI-safe demonstration on synthetic data:

DOI list
    │
    ▼
[Fetch]      retrieve bibliographic metadata for each DOI
    │
    ▼
[Enrich]     add structured AI summary, keywords, and topic tags
    │
    ▼
[Validate]   enforce schema on every enriched record
    │
    ▼
[Output]     per-record JSON  +  manifest with SHA-256 checksums

One command turns a list of identifiers into validated, checksummed records ready for downstream ingest — no manual extraction or formatting in between. It runs out of the box with no API keys (the default mock provider is deterministic and CI-safe), so you can verify the architecture and verification pattern in seconds.


Why You Can Trust the Output

This is not a thin wrapper around an LLM. The architecture is designed so the output can be relied on:

  • Reproducible — deterministic processing; the same input produces the same output.
  • Auditable — every run produces a manifest with per-record SHA-256 checksums.
  • Validated — schema enforcement on each record; malformed output fails fast.
  • Provider-agnostic — metadata and LLM backends swap behind a stable interface, without touching pipeline logic.

These are the same properties — applied here to the enrichment layer — that make the production system above trustworthy.


Quick Start

git clone https://github.com/DmitryIri/ai-editorial-digest-pipeline.git
cd ai-editorial-digest-pipeline
pip install -e ".[dev]"
python examples/quickstart.py

Expected output:

Enriched 2 records:

DOI:        10.9999/synthetic.2024.001
Title:      Automated Knowledge Extraction from Scientific Literature
AI Summary: This paper titled 'Automated Knowledge Extraction from
            Scientific Literature' presents research in the domain of
            NLP, knowledge extraction.
AI Topics:  methodology, empirical study
Provider:   mock

Manifest root checksum: 68f60e9684391180...
Records:                2

Runs out of the box with no API keys — the default mock provider is CI-safe and deterministic.


Output Format

Each enriched record:

{
  "doi": "10.9999/synthetic.2024.001",
  "title": "Automated Knowledge Extraction from Scientific Literature",
  "abstract": "We present a synthetic study on automated extraction of structured knowledge...",
  "authors": ["Alice Researcher", "Bob Scholar"],
  "year": 2024,
  "journal": "Journal of Synthetic Research",
  "keywords": ["NLP", "knowledge extraction"],
  "ai_summary": "This paper titled 'Automated Knowledge Extraction from Scientific Literature' presents research in the domain of NLP, knowledge extraction.",
  "ai_keywords": ["ai-generated", "mock", "NLP", "knowledge extraction", "scientific literature"],
  "ai_topics": ["methodology", "empirical study"],
  "enrichment_provider": "mock"
}

Manifest with reproducibility checksums:

{
  "generated_at": "2024-01-15T10:30:00+00:00",
  "record_count": 2,
  "items": [{"doi": "10.9999/synthetic.2024.001", "checksum": "sha256..."}],
  "root_checksum": "sha256..."
}

Architecture (public repo)

DOI list
    │
    ▼
[Fetcher]           ← METADATA_PROVIDER=mock | crossref
    │ RawMetadata (Pydantic)
    ▼
[Enricher]          ← LLM_PROVIDER=mock | openai | anthropic
    │ EnrichedMetadata (Pydantic)
    ▼
[Validator]         ← schema enforcement
    │
    ▼
JSON records + Manifest (SHA-256 checksums)

A clean separation of concerns: each stage has a typed contract, so backends can change without rewriting the pipeline.


Potential Business Applications

The same pattern — take an external content stream, turn it into validated, traceable, standardized output — applies wherever expert time is spent on repeatable extraction and structuring:

  • Research portals — auto-enrich imported bibliographic records
  • Publishing automation — metadata completion and quality checks before ingest
  • Literature review / RAG preparation — structured extraction for downstream pipelines
  • Knowledge graph population — structured entities from abstracts at scale

Running Tests

pytest tests/ -v
ruff check src/ tests/ examples/ tools/

All tests run on the mock provider — no API keys required. CI runs on Python 3.10 / 3.11 / 3.12.


Current Scope

This public repository is the enrichment layer, packaged as a runnable, CI-safe demonstration:

  • Default mode is mock — synthetic samples, no API keys, deterministic output. Safe for CI and for running in seconds.
  • Live providers are optional integration points — METADATA_PROVIDER=crossref and LLM_PROVIDER=openai|anthropic are defined interfaces ready to be wired to real backends.
  • Sample data in fixtures/ is synthetic — no real publications or authors.

The architecture, contracts, validation, and reproducibility pattern are real; the live API backends are the layer left open for adaptation to a specific client's sources and standards. The translation/editing and document-generation stages of the production system are domain- and template-specific and are not part of this public demonstration.


Stack (public repo)

  • Python 3.10+
  • Pydantic v2 — data validation and schema enforcement
  • httpx — HTTP client for metadata integration
  • ruff — linting · pytest + pytest-cov — testing
  • GitHub Actions CI (Python 3.10 / 3.11 / 3.12)

License

MIT

About

Validated, audit-ready metadata enrichment for research & publishing workflows — reproducible AI pipeline patterned on a production editorial system.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages