Skip to content

Repository files navigation

Infoscience Import Pipeline

License: MIT Python 3.11+

Automated harvesting, deduplication, enrichment, and loading of publication data from multiple external sources into Infoscience / DSpace-CRIS.

Sources: Scopus · Web of Science · Crossref · OpenAlex · Zenodo · EPO OPS


Table of contents

  1. What it does
  2. Prerequisites
  3. Installation
  4. Quick start
  5. Streamlit supervision UI
  6. Authentication
  7. Scheduled runs
  8. CLI reference
  9. Environments
  10. Environment variables
  11. Architecture
  12. Output structure
  13. Incremental logic
  14. License
  15. Citation

What it does

The pipeline runs a linear sequence of stages on every execution:

Stage Description
Harvest Queries each enabled source API within the configured time window
Deduplicate Cross-source dedup (DOI, then title + year with type-aware rules for preprints and datasets), then type-scoped dedup against existing Infoscience items; ambiguous cases are flagged and forwarded to the DSpace workspace
Enrich EPFL author reconciliation (People API, ORCID); OA/full-text metadata (Unpaywall, OpenAlex)
Load Builds DSpace-CRIS item payloads and ingests them (skipped in --dry-run)
Report Generates a timestamped Excel report and optionally sends it by email
Persist Writes run history, per-source stats, publications, and EPFL author data to DuckDB

The pipeline is stateless: incrementality comes from the sliding time window and deduplication against what is already in Infoscience.


Prerequisites

  • Python 3.11+
  • A .env.dev / .env.test / .env.prod credentials file at the project root (see Environment variables)
  • Network access to the APIs you intend to harvest

Installation

git clone <repo-url> && cd infoscience-imports
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

# Create your first credentials file
cp .sample.env .env.dev   # edit and fill in values

Quick start

# Dry-run on dev — harvests and enriches but does not load into DSpace
python3 data_pipeline/main.py --env dev --dry-run --no-email -vv

# Real run on dev with a 7-day window
python3 data_pipeline/main.py --env dev --window-days 7

# Production run
python3 data_pipeline/main.py --env prod

Streamlit supervision UI

A web dashboard for monitoring run history, browsing and curating publications, and launching pipeline runs without touching the CLI.

./run_ui.sh          # default port 8501
./run_ui.sh 8502     # custom port
# or directly:
streamlit run app.py

Pages

Page Role Description
🏠 Tableau de bord all KPIs, 30-day trend chart, recent runs, per-source breakdown, charts by type / OA status / year / unit / journal / PDF proportion
🚀 Lancer un run admin Form to configure and launch the pipeline; live log streaming with stop button
⏰ Programmation admin Create and manage scheduled runs; cron-based triggers; enable/disable toggle; run-now button; scheduler status indicator
📋 Publications all Paginated, sortable, filterable datatable; OA / licence / PDF badges; EPFL author + unit aggregation; weak-status flag; Excel report download
📊 Statistiques all Per-run funnel by source, publication type breakdown, EPFL author and unit tabs
⚙️ Configuration admin Environment variable status, DuckDB info, .env template

Dashboard charts

The dashboard includes six analytical visualisations, all scoped to the selected run or global:

  • Types de documents — horizontal bar, top-15
  • Statut Open Access — donut (OA + PDF / OA sans PDF / OA non-libre / Non-OA / Non défini)
  • Publications par année — bar chart by publication year
  • Proportion PDF récupéré — donut with summary count
  • Top unités EPFL — horizontal bar, top-15
  • Top journaux — horizontal bar, top-15

Publication curation features

The Publications page supports detailed curation:

  • OA columnOA / Non-OA / Non-libre (publisher-specific licences are never treated as OA)
  • Licence column — CC-BY variant displayed (e.g. CC-BY-NC-ND); prefix-matched so all cc-* variants are covered
  • PDF ✓ column — checkbox; only True when a valid PDF was retrieved under an open licence
  • Auteurs EPFL — reconciled EPFL authors with status/position; for rejected publications, shows pre-detected unreconciled authors
  • Unités — EPFL units with type in parentheses
  • ⚠️ column — flags publications where all matched EPFL authors have a "weak" status (Hôte, Hors EPFL, Étudiant, or Personnel with non-permanent position)
  • Note dédup (dedup_note) — set when a record was let through despite a potential conflict in Infoscience; possible values: supersedes_preprint, published_version_exists, cross_type_doi, dataset_in_other_collection
  • 🚩 Doublon Infoscience (flagged_publication) — JSON list of existing Infoscience items that triggered the flag, each with uuid, doi, and dc_type; use the UUID to locate the item directly in Infoscience

Use the Signalement dédup filter (options: Tous / 🚩 Flaggés / specific note value) to isolate flagged records for curation. Flagged records also appear in the dedicated Flagged Publications sheet of the Excel report.

The environment selector in the sidebar switches between dev, test, and prod — the choice is persisted and automatically passed to any run launched from the UI.


Authentication

The Streamlit UI requires login. Two roles are available:

Role Pages
admin All pages
reporting Tableau de bord · Publications · Statistiques

The CLI pipeline is not protected by authentication — credentials are only required for the web UI.

Setup

# Create the first users (passwords are prompted interactively)
python -m ui.auth add admin    admin
python -m ui.auth add reporter reporting

# Other management commands
python -m ui.auth list
python -m ui.auth passwd <username>
python -m ui.auth remove <username>

Credentials are stored as bcrypt hashes in .streamlit/auth.yaml (gitignored). Copy .streamlit/auth.yaml.example as a reference for the file structure.


Scheduled runs

The ⏰ Programmation page (admin only) lets you define pipeline runs that fire automatically on a cron schedule, without manual intervention.

How it works

run_ui.sh starts scheduler.py as a background daemon alongside Streamlit. The daemon reads data/schedules.json every 15 seconds and registers or updates APScheduler jobs accordingly — any change made in the UI takes effect without a restart.

When a scheduled job fires, the scheduler:

  1. Acquires the per-environment run lock (data/run_active_{env}.json). If a run is already active, the execution is silently skipped.
  2. Spawns data_pipeline/main.py as an independent OS subprocess.
  3. Updates last_run_at, last_run_id, and last_run_status in schedules.json on completion.

Because jobs run as separate OS processes, stopping or restarting the UI does not interrupt a run already in progress — it continues to completion and releases the lock itself.

Creating a schedule

  1. Open ⏰ Programmation in the sidebar (admin role required).
  2. Check the scheduler status indicator at the top — it should show 🟢. If it shows ⚪, start the UI via ./run_ui.sh.
  3. Click ➕ Nouveau schedule and fill in:
Field Description
Nom Label shown in the schedule list
Environnement dev, test, or prod
Sources One or more sources to harvest (default: scopus, crossref, openalex)
Fenêtre glissante Days back from today (default: 20)
Fréquence Preset (daily, weekly, …) or "Personnalisé…" for a manual cron expression
Expression cron Standard 5-field cron — pre-filled from the preset, always editable
Dry-run Skip DSpace ingestion
Désactiver l'envoi d'e-mail Suppress the Excel report email
  1. Click Créer le schedule. The next execution time is shown immediately.

Managing schedules

Each schedule card shows the next scheduled execution and the last run result (✅ completed / ⏳ running / ❌ failed / 🛑 killed).

  • Toggle "Actif" — enable or disable without deleting the schedule. Takes effect within 15 seconds.
  • ▶ Now — fire the run immediately, using the schedule's configuration.
  • 🗑 — permanently delete the schedule.

The last 50 lines of logs/scheduler.log are accessible at the bottom of the page.


CLI reference

Time window

Argument Default Description
--window-days N 15 Sliding window: today − N daystoday
--start-date YYYY-MM-DD Fixed start date (requires --end-date)
--end-date YYYY-MM-DD Fixed end date (requires --start-date)

Sources

# All sources (default)
python3 data_pipeline/main.py

# Specific subset
python3 data_pipeline/main.py --sources scopus,wos,openalex

Available: wos, scopus, crossref, openalex, zenodo, epo

Query overrides

Override the default institution-wide query for one or more sources:

python3 data_pipeline/main.py \
  --query-wos    "EPFL OR Lausanne" \
  --query-scopus "AFFIL(EPFL)"

Query overrides are also available in the UI via the Requêtes (optionnel) expander on the run page.

Crossref — flexible query format

The Crossref query field accepts three formats:

Plain string — uses the generic query parameter:

--query-crossref "EPFL machine learning"

JSON object — spread directly as API parameters (supports any Crossref query index or filter):

--query-crossref '{"query.affiliation": "EPFL SV", "filter": "type:journal-article"}'
--query-crossref '{"filter": "orcid:0000-0002-1825-0097"}'

JSON array — each element runs as a separate API call; results are merged and deduplicated by DOI:

--query-crossref '[{"query.affiliation": "EPFL"}, {"filter": "ror-id:02s376052"}]'

Reserved filters: from-created-date and until-created-date are always injected by the harvester (from the configured time window) and cannot be overridden. Any other Crossref filter can be used freely.

Author ID-based harvesting

When any of these flags are provided, the pipeline switches to ID-based harvesting for the relevant sources instead of institution-wide queries.

# Scopus Author IDs
python3 data_pipeline/main.py --scopus-ids "7004212771,57201854951"

# Web of Science ResearcherIDs
python3 data_pipeline/main.py --wos-ids "A-1234-2010,R-5678-2017"

# ORCID iDs (used by Crossref and OpenAlex)
python3 data_pipeline/main.py --orcid-ids "0000-0002-1825-0097"

# OpenAlex Author IDs
python3 data_pipeline/main.py --openalex-ids "A12345678"

# All combined
python3 data_pipeline/main.py \
  --scopus-ids "7004212771" \
  --orcid-ids  "0000-0002-1825-0097" \
  --sources scopus,openalex,crossref

Each flag also accepts a file path (one ID per line):

python3 data_pipeline/main.py --scopus-ids ./ids/scopus_ids.txt

Run modes

Flag Effect
--dry-run Skip DSpace load and email — safe for inspection
--no-email Generate the report but do not send it
-v / -vv Increase log verbosity
--env {dev,test,prod} Select the target environment (see Environments)
--output-dir PATH Override the output directory (default: data/)

Common recipes

# Inspect last 7 days without touching anything
python3 data_pipeline/main.py --dry-run --no-email --window-days 7 -vv

# Backfill a specific month on prod
python3 data_pipeline/main.py --env prod \
  --start-date 2025-01-01 --end-date 2025-01-31

# Harvest one author across all relevant sources
python3 data_pipeline/main.py \
  --scopus-ids "7004212771" \
  --orcid-ids  "0000-0002-1825-0097" \
  --wos-ids    "A-1234-2010" \
  --sources scopus,wos,openalex,crossref --dry-run

Environments (dev / test / prod)

The pipeline supports three fully isolated environments. Each has its own credentials file, DuckDB database, and run-lock — a production run never interferes with development.

Environment Credentials Database Run lock
dev .env.dev data/pipeline_dev.duckdb data/run_active_dev.json
test .env.test data/pipeline_test.duckdb data/run_active_test.json
prod .env.prod data/pipeline_prod.duckdb data/run_active_prod.json

The active environment defaults to dev. All .env.* files are git-ignored.

Setup

cp .sample.env .env.dev   # fill in dev credentials
cp .sample.env .env.test  # fill in test credentials
cp .sample.env .env.prod  # fill in prod credentials

Switching environments

CLI — applies to that invocation only, not persisted:

python3 data_pipeline/main.py --env prod --dry-run

Streamlit UI — use the selector at the top of the sidebar; the choice is persisted to data/active_env and passed automatically to any run launched from the UI. A coloured badge (green / orange / red) and a warning banner indicate the active environment.

Shell variable — highest priority, useful for cron or CI/CD:

APP_ENV=prod python3 data_pipeline/main.py
APP_ENV=test streamlit run app.py

Priority order: APP_ENV > --env > persisted data/active_env > dev

Fallback: if .env.{env} does not exist, the pipeline falls back to the generic .env at the project root.


Environment variables

Copy .sample.env to the appropriate .env.* file(s) and fill in the values.

Required

Variable Description
DS_API_ENDPOINT DSpace REST API base URL — https://<domain>/server/api
DS_API_TOKEN DSpace REST API static token

Authentication (alternative)

Variable Description
DS_ACCESS_TOKEN DSpace session cookie token — alternative auth set after login

EPFL People API

Variable Description
API_EPFL_USER EPFL People API username (author reconciliation)
API_EPFL_PWD EPFL People API password

Harvesting sources

Variable Source Description
SCOPUS_API_KEY Scopus Elsevier API key
SCOPUS_INST_TOKEN Scopus Elsevier institutional token
ELS_API_KEY Unpaywall Elsevier key for full-text PDF retrieval
WOS_TOKEN Web of Science Clarivate API token
EPO_OPS_KEY EPO OPS Open Patent Services consumer key
EPO_OPS_SECRET EPO OPS Open Patent Services consumer secret
OPENALEX_API_KEY OpenAlex Authenticated access (higher rate limits)
OPENALEX_DATA_VERSION OpenAlex API data version (default: 2)
ZENODO_API_KEY Zenodo Authenticated rate limit
ORCID_API_TOKEN ORCID Bearer token for author reconciliation

Polite pool / HTTP

Variable Description
CONTACT_API_EMAIL Email sent as mailto in requests to Crossref, Unpaywall, OpenAlex — strongly recommended
USER_AGENT HTTP User-Agent header (defaults to a sensible EPFL string if unset)

Email report (optional)

Variable Description
RECIPIENT_EMAIL Report recipient address
SENDER_EMAIL Report sender address
SMTP_SERVER SMTP hostname

Architecture

The pipeline follows a strict linear sequence: Harvest → Deduplicate → Enrich → Load → Report → Persist.

data_pipeline/
├── main.py          Entry point — CLI args, orchestration, DuckDB persistence
├── harvester.py     One Harvester subclass per source (WoS, Scopus, Crossref, …)
├── deduplicator.py  Cross-source dedup + DSpace-aware dedup
├── enricher.py      EPFL author reconciliation, OA/full-text enrichment
├── loader.py        Builds DSpace-CRIS payloads, calls DSpaceClientWrapper
└── reporting.py     Excel report generation + SMTP delivery

clients/             One module per external API
config/
├── __init__.py      YAML loader — exposes source_order, default_queries, unit_types, …
├── pipeline.yaml    Default harvest queries, source priority order, unit filters, Scopus AF-IDs
└── mappings/
    ├── collections.yaml      Infoscience collection names → UUIDs + DSpace Submission Forms section names
    ├── doctypes.yaml         Source doc-types → collection + dc.type (active + commented-out)
    ├── licenses.yaml         OA licence identifiers → DSpace display values
    ├── versions.yaml         OA version identifiers → COAR URIs
    └── types_authority.yaml  dc.type values → COAR authority identifiers
mappings.py          Loads the above YAML files; exposes classify_record_type, get_version_mapping, …
env_loader.py        Environment selection and .env.* loading
db/pipeline_db.py    DuckDB persistence layer (run history, publications, authors)
ui/
├── run_state.py     File-based mutex (one pipeline run at a time, per environment)
├── auth.py          Streamlit authentication + role-based ACL
└── styles.css       External stylesheet (colour tokens injected from app.py)
app.py               Streamlit supervision UI

Key design points:

  • env_loader.py is the single source of truth for environment selection. It sets APP_ENV in os.environ on load so every downstream module (PipelineDB, run_state) picks up the correct environment without re-reading the state file.
  • PipelineDB opens a new DuckDB connection for every operation and closes it immediately — no persistent connections, which avoids write-lock conflicts. All DDL uses CREATE/ALTER … IF NOT EXISTS so schema migrations are idempotent.
  • The run-lock file (data/run_active_{env}.json) is created atomically with open(..., 'x') so two simultaneous UI submissions cannot both start a run.
  • The UI is authenticated via bcrypt-hashed credentials in .streamlit/auth.yaml. The CLI is authentication-free.
  • All data-driven configuration (queries, mappings, collection UUIDs) lives in config/pipeline.yaml and config/mappings/*.yaml. To add a new document type, update doctypes.yaml; to update a collection UUID after a DSpace migration, update collections.yaml — no Python changes required.
  • Source priority for deduplication merging is defined in config/pipeline.yaml → source_order.
  • The stylesheet lives in ui/styles.css (pure CSS); app.py only injects colour tokens as CSS custom properties (var(--canard), etc.) via a small inline <style> block.

Output structure

Each run produces a timestamped subfolder under data/:

data/
└── 2025-10-08_03-15/
    ├── Raw_WosItems.csv
    ├── Raw_ScopusItems.csv
    ├── Raw_CrossrefItems.csv
    ├── Raw_OpenalexItems.csv
    ├── Raw_ZenodoItems.csv
    ├── Raw_EpoItems.csv
    ├── DeduplicatedItems.csv
    ├── UnloadedItems.csv          # clear duplicates found in Infoscience (discarded)
    ├── Items.csv
    ├── AuthorsAndAffiliations.csv
    ├── EpflAuthors.csv
    ├── ItemsWithOAMetadata.csv
    ├── ImportedItems.csv
    ├── RejectedItems.csv
    └── Report_2025-10-08_03-15.xlsx

Run history, per-source statistics, publications, EPFL authors, and unit links are also written to DuckDB (data/pipeline_{env}.duckdb) and are browsable through the Streamlit UI.


Incremental logic

The pipeline is intentionally stateless. Re-running it is always safe because:

  1. Sliding window — only publications within the configured date range are harvested.
  2. DSpace-aware dedup — the deduplicator queries Infoscience for existing items before loading, so already-imported records are never duplicated. Deduplication is type-aware:
    • A dataset is only deduplicated against other datasets (scoped to the Datasets and Code collection); a title match in another collection is flagged but not discarded.
    • A preprint whose published version already exists in Infoscience is forwarded to the DSpace workspace rather than silently dropped, so it can be reviewed and linked.
    • Clear duplicates (same DOI, same type, same title + year within the same collection) are discarded without flagging.
  3. Stable row_id — each record gets a deterministic hash of its key fields, ensuring consistent matching across runs.

Running daily with a 15-day window (the default) catches late-indexed publications while the overlap with previous windows is handled entirely by deduplication.


License

This project is licensed under the MIT License — see LICENSE for the full text.

The bundled DSpace REST Python Client (dspace/dspace_rest_client/) is licensed under the BSD 3-Clause License (© The Library Code GmbH). MIT and BSD-3-Clause are compatible permissive licences; the BSD-3-Clause copyright notice is preserved in the LICENSE file and in the bundled source.


Citation

If you use this software in your research or institutional work, please cite it using the metadata in CITATION.cff or the reference below:

@software{infoscience_import_pipeline,
  author    = {Sicot, Julien and Borel, Alain and Geoffroy, Géraldine},
  title     = {Infoscience Import Pipeline},
  year      = {2026},
  publisher = {EPFL Library},
  url       = {https://github.com/epfllibrary/infoscience-imports},
  license   = {MIT}
}

About

Infoscience Imports uses modular Python scripts for distinct tasks: harvesting publications and files from different sources : OpenAlex, unpaywall, WoS and Scopus, deduplicating and enriching metadata via local API, and uploading records to Infoscience using the DSpace API.

Topics

Resources

Stars

6 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages