Automated harvesting, deduplication, enrichment, and loading of publication data from multiple external sources into Infoscience / DSpace-CRIS.
Sources: Scopus · Web of Science · Crossref · OpenAlex · Zenodo · EPO OPS
- What it does
- Prerequisites
- Installation
- Quick start
- Streamlit supervision UI
- Authentication
- Scheduled runs
- CLI reference
- Environments
- Environment variables
- Architecture
- Output structure
- Incremental logic
- License
- Citation
The pipeline runs a linear sequence of stages on every execution:
| Stage | Description |
|---|---|
| Harvest | Queries each enabled source API within the configured time window |
| Deduplicate | Cross-source dedup (DOI, then title + year with type-aware rules for preprints and datasets), then type-scoped dedup against existing Infoscience items; ambiguous cases are flagged and forwarded to the DSpace workspace |
| Enrich | EPFL author reconciliation (People API, ORCID); OA/full-text metadata (Unpaywall, OpenAlex) |
| Load | Builds DSpace-CRIS item payloads and ingests them (skipped in --dry-run) |
| Report | Generates a timestamped Excel report and optionally sends it by email |
| Persist | Writes run history, per-source stats, publications, and EPFL author data to DuckDB |
The pipeline is stateless: incrementality comes from the sliding time window and deduplication against what is already in Infoscience.
- Python 3.11+
- A
.env.dev/.env.test/.env.prodcredentials file at the project root (see Environment variables) - Network access to the APIs you intend to harvest
git clone <repo-url> && cd infoscience-imports
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Create your first credentials file
cp .sample.env .env.dev # edit and fill in values# Dry-run on dev — harvests and enriches but does not load into DSpace
python3 data_pipeline/main.py --env dev --dry-run --no-email -vv
# Real run on dev with a 7-day window
python3 data_pipeline/main.py --env dev --window-days 7
# Production run
python3 data_pipeline/main.py --env prodA web dashboard for monitoring run history, browsing and curating publications, and launching pipeline runs without touching the CLI.
./run_ui.sh # default port 8501
./run_ui.sh 8502 # custom port
# or directly:
streamlit run app.py| Page | Role | Description |
|---|---|---|
| 🏠 Tableau de bord | all | KPIs, 30-day trend chart, recent runs, per-source breakdown, charts by type / OA status / year / unit / journal / PDF proportion |
| 🚀 Lancer un run | admin | Form to configure and launch the pipeline; live log streaming with stop button |
| ⏰ Programmation | admin | Create and manage scheduled runs; cron-based triggers; enable/disable toggle; run-now button; scheduler status indicator |
| 📋 Publications | all | Paginated, sortable, filterable datatable; OA / licence / PDF badges; EPFL author + unit aggregation; weak-status flag; Excel report download |
| 📊 Statistiques | all | Per-run funnel by source, publication type breakdown, EPFL author and unit tabs |
| ⚙️ Configuration | admin | Environment variable status, DuckDB info, .env template |
The dashboard includes six analytical visualisations, all scoped to the selected run or global:
- Types de documents — horizontal bar, top-15
- Statut Open Access — donut (OA + PDF / OA sans PDF / OA non-libre / Non-OA / Non défini)
- Publications par année — bar chart by publication year
- Proportion PDF récupéré — donut with summary count
- Top unités EPFL — horizontal bar, top-15
- Top journaux — horizontal bar, top-15
The Publications page supports detailed curation:
- OA column —
OA/Non-OA/Non-libre(publisher-specific licences are never treated as OA) - Licence column — CC-BY variant displayed (e.g.
CC-BY-NC-ND); prefix-matched so allcc-*variants are covered - PDF ✓ column — checkbox; only
Truewhen a valid PDF was retrieved under an open licence - Auteurs EPFL — reconciled EPFL authors with status/position; for rejected publications, shows pre-detected unreconciled authors
- Unités — EPFL units with type in parentheses
⚠️ column — flags publications where all matched EPFL authors have a "weak" status (Hôte, Hors EPFL, Étudiant, or Personnel with non-permanent position)- Note dédup (
dedup_note) — set when a record was let through despite a potential conflict in Infoscience; possible values:supersedes_preprint,published_version_exists,cross_type_doi,dataset_in_other_collection - 🚩 Doublon Infoscience (
flagged_publication) — JSON list of existing Infoscience items that triggered the flag, each withuuid,doi, anddc_type; use the UUID to locate the item directly in Infoscience
Use the Signalement dédup filter (options: Tous / 🚩 Flaggés / specific note value) to isolate flagged records for curation. Flagged records also appear in the dedicated Flagged Publications sheet of the Excel report.
The environment selector in the sidebar switches between dev, test, and prod — the choice is persisted and automatically passed to any run launched from the UI.
The Streamlit UI requires login. Two roles are available:
| Role | Pages |
|---|---|
admin |
All pages |
reporting |
Tableau de bord · Publications · Statistiques |
The CLI pipeline is not protected by authentication — credentials are only required for the web UI.
# Create the first users (passwords are prompted interactively)
python -m ui.auth add admin admin
python -m ui.auth add reporter reporting
# Other management commands
python -m ui.auth list
python -m ui.auth passwd <username>
python -m ui.auth remove <username>Credentials are stored as bcrypt hashes in .streamlit/auth.yaml (gitignored). Copy .streamlit/auth.yaml.example as a reference for the file structure.
The ⏰ Programmation page (admin only) lets you define pipeline runs that fire automatically on a cron schedule, without manual intervention.
run_ui.sh starts scheduler.py as a background daemon alongside Streamlit. The daemon reads data/schedules.json every 15 seconds and registers or updates APScheduler jobs accordingly — any change made in the UI takes effect without a restart.
When a scheduled job fires, the scheduler:
- Acquires the per-environment run lock (
data/run_active_{env}.json). If a run is already active, the execution is silently skipped. - Spawns
data_pipeline/main.pyas an independent OS subprocess. - Updates
last_run_at,last_run_id, andlast_run_statusinschedules.jsonon completion.
Because jobs run as separate OS processes, stopping or restarting the UI does not interrupt a run already in progress — it continues to completion and releases the lock itself.
- Open ⏰ Programmation in the sidebar (admin role required).
- Check the scheduler status indicator at the top — it should show 🟢. If it shows ⚪, start the UI via
./run_ui.sh. - Click ➕ Nouveau schedule and fill in:
| Field | Description |
|---|---|
| Nom | Label shown in the schedule list |
| Environnement | dev, test, or prod |
| Sources | One or more sources to harvest (default: scopus, crossref, openalex) |
| Fenêtre glissante | Days back from today (default: 20) |
| Fréquence | Preset (daily, weekly, …) or "Personnalisé…" for a manual cron expression |
| Expression cron | Standard 5-field cron — pre-filled from the preset, always editable |
| Dry-run | Skip DSpace ingestion |
| Désactiver l'envoi d'e-mail | Suppress the Excel report email |
- Click Créer le schedule. The next execution time is shown immediately.
Each schedule card shows the next scheduled execution and the last run result (✅ completed / ⏳ running / ❌ failed / 🛑 killed).
- Toggle "Actif" — enable or disable without deleting the schedule. Takes effect within 15 seconds.
- ▶ Now — fire the run immediately, using the schedule's configuration.
- 🗑 — permanently delete the schedule.
The last 50 lines of logs/scheduler.log are accessible at the bottom of the page.
| Argument | Default | Description |
|---|---|---|
--window-days N |
15 |
Sliding window: today − N days → today |
--start-date YYYY-MM-DD |
— | Fixed start date (requires --end-date) |
--end-date YYYY-MM-DD |
— | Fixed end date (requires --start-date) |
# All sources (default)
python3 data_pipeline/main.py
# Specific subset
python3 data_pipeline/main.py --sources scopus,wos,openalexAvailable: wos, scopus, crossref, openalex, zenodo, epo
Override the default institution-wide query for one or more sources:
python3 data_pipeline/main.py \
--query-wos "EPFL OR Lausanne" \
--query-scopus "AFFIL(EPFL)"Query overrides are also available in the UI via the Requêtes (optionnel) expander on the run page.
The Crossref query field accepts three formats:
Plain string — uses the generic query parameter:
--query-crossref "EPFL machine learning"JSON object — spread directly as API parameters (supports any Crossref query index or filter):
--query-crossref '{"query.affiliation": "EPFL SV", "filter": "type:journal-article"}'
--query-crossref '{"filter": "orcid:0000-0002-1825-0097"}'JSON array — each element runs as a separate API call; results are merged and deduplicated by DOI:
--query-crossref '[{"query.affiliation": "EPFL"}, {"filter": "ror-id:02s376052"}]'Reserved filters:
from-created-dateanduntil-created-dateare always injected by the harvester (from the configured time window) and cannot be overridden. Any other Crossref filter can be used freely.
When any of these flags are provided, the pipeline switches to ID-based harvesting for the relevant sources instead of institution-wide queries.
# Scopus Author IDs
python3 data_pipeline/main.py --scopus-ids "7004212771,57201854951"
# Web of Science ResearcherIDs
python3 data_pipeline/main.py --wos-ids "A-1234-2010,R-5678-2017"
# ORCID iDs (used by Crossref and OpenAlex)
python3 data_pipeline/main.py --orcid-ids "0000-0002-1825-0097"
# OpenAlex Author IDs
python3 data_pipeline/main.py --openalex-ids "A12345678"
# All combined
python3 data_pipeline/main.py \
--scopus-ids "7004212771" \
--orcid-ids "0000-0002-1825-0097" \
--sources scopus,openalex,crossrefEach flag also accepts a file path (one ID per line):
python3 data_pipeline/main.py --scopus-ids ./ids/scopus_ids.txt| Flag | Effect |
|---|---|
--dry-run |
Skip DSpace load and email — safe for inspection |
--no-email |
Generate the report but do not send it |
-v / -vv |
Increase log verbosity |
--env {dev,test,prod} |
Select the target environment (see Environments) |
--output-dir PATH |
Override the output directory (default: data/) |
# Inspect last 7 days without touching anything
python3 data_pipeline/main.py --dry-run --no-email --window-days 7 -vv
# Backfill a specific month on prod
python3 data_pipeline/main.py --env prod \
--start-date 2025-01-01 --end-date 2025-01-31
# Harvest one author across all relevant sources
python3 data_pipeline/main.py \
--scopus-ids "7004212771" \
--orcid-ids "0000-0002-1825-0097" \
--wos-ids "A-1234-2010" \
--sources scopus,wos,openalex,crossref --dry-runThe pipeline supports three fully isolated environments. Each has its own credentials file, DuckDB database, and run-lock — a production run never interferes with development.
| Environment | Credentials | Database | Run lock |
|---|---|---|---|
dev |
.env.dev |
data/pipeline_dev.duckdb |
data/run_active_dev.json |
test |
.env.test |
data/pipeline_test.duckdb |
data/run_active_test.json |
prod |
.env.prod |
data/pipeline_prod.duckdb |
data/run_active_prod.json |
The active environment defaults to dev. All .env.* files are git-ignored.
cp .sample.env .env.dev # fill in dev credentials
cp .sample.env .env.test # fill in test credentials
cp .sample.env .env.prod # fill in prod credentialsCLI — applies to that invocation only, not persisted:
python3 data_pipeline/main.py --env prod --dry-runStreamlit UI — use the selector at the top of the sidebar; the choice is persisted to data/active_env and passed automatically to any run launched from the UI. A coloured badge (green / orange / red) and a warning banner indicate the active environment.
Shell variable — highest priority, useful for cron or CI/CD:
APP_ENV=prod python3 data_pipeline/main.py
APP_ENV=test streamlit run app.pyPriority order: APP_ENV > --env > persisted data/active_env > dev
Fallback: if .env.{env} does not exist, the pipeline falls back to the generic .env at the project root.
Copy .sample.env to the appropriate .env.* file(s) and fill in the values.
| Variable | Description |
|---|---|
DS_API_ENDPOINT |
DSpace REST API base URL — https://<domain>/server/api |
DS_API_TOKEN |
DSpace REST API static token |
| Variable | Description |
|---|---|
DS_ACCESS_TOKEN |
DSpace session cookie token — alternative auth set after login |
| Variable | Description |
|---|---|
API_EPFL_USER |
EPFL People API username (author reconciliation) |
API_EPFL_PWD |
EPFL People API password |
| Variable | Source | Description |
|---|---|---|
SCOPUS_API_KEY |
Scopus | Elsevier API key |
SCOPUS_INST_TOKEN |
Scopus | Elsevier institutional token |
ELS_API_KEY |
Unpaywall | Elsevier key for full-text PDF retrieval |
WOS_TOKEN |
Web of Science | Clarivate API token |
EPO_OPS_KEY |
EPO OPS | Open Patent Services consumer key |
EPO_OPS_SECRET |
EPO OPS | Open Patent Services consumer secret |
OPENALEX_API_KEY |
OpenAlex | Authenticated access (higher rate limits) |
OPENALEX_DATA_VERSION |
OpenAlex | API data version (default: 2) |
ZENODO_API_KEY |
Zenodo | Authenticated rate limit |
ORCID_API_TOKEN |
ORCID | Bearer token for author reconciliation |
| Variable | Description |
|---|---|
CONTACT_API_EMAIL |
Email sent as mailto in requests to Crossref, Unpaywall, OpenAlex — strongly recommended |
USER_AGENT |
HTTP User-Agent header (defaults to a sensible EPFL string if unset) |
| Variable | Description |
|---|---|
RECIPIENT_EMAIL |
Report recipient address |
SENDER_EMAIL |
Report sender address |
SMTP_SERVER |
SMTP hostname |
The pipeline follows a strict linear sequence: Harvest → Deduplicate → Enrich → Load → Report → Persist.
data_pipeline/
├── main.py Entry point — CLI args, orchestration, DuckDB persistence
├── harvester.py One Harvester subclass per source (WoS, Scopus, Crossref, …)
├── deduplicator.py Cross-source dedup + DSpace-aware dedup
├── enricher.py EPFL author reconciliation, OA/full-text enrichment
├── loader.py Builds DSpace-CRIS payloads, calls DSpaceClientWrapper
└── reporting.py Excel report generation + SMTP delivery
clients/ One module per external API
config/
├── __init__.py YAML loader — exposes source_order, default_queries, unit_types, …
├── pipeline.yaml Default harvest queries, source priority order, unit filters, Scopus AF-IDs
└── mappings/
├── collections.yaml Infoscience collection names → UUIDs + DSpace Submission Forms section names
├── doctypes.yaml Source doc-types → collection + dc.type (active + commented-out)
├── licenses.yaml OA licence identifiers → DSpace display values
├── versions.yaml OA version identifiers → COAR URIs
└── types_authority.yaml dc.type values → COAR authority identifiers
mappings.py Loads the above YAML files; exposes classify_record_type, get_version_mapping, …
env_loader.py Environment selection and .env.* loading
db/pipeline_db.py DuckDB persistence layer (run history, publications, authors)
ui/
├── run_state.py File-based mutex (one pipeline run at a time, per environment)
├── auth.py Streamlit authentication + role-based ACL
└── styles.css External stylesheet (colour tokens injected from app.py)
app.py Streamlit supervision UI
Key design points:
env_loader.pyis the single source of truth for environment selection. It setsAPP_ENVinos.environon load so every downstream module (PipelineDB,run_state) picks up the correct environment without re-reading the state file.PipelineDBopens a new DuckDB connection for every operation and closes it immediately — no persistent connections, which avoids write-lock conflicts. All DDL usesCREATE/ALTER … IF NOT EXISTSso schema migrations are idempotent.- The run-lock file (
data/run_active_{env}.json) is created atomically withopen(..., 'x')so two simultaneous UI submissions cannot both start a run. - The UI is authenticated via bcrypt-hashed credentials in
.streamlit/auth.yaml. The CLI is authentication-free. - All data-driven configuration (queries, mappings, collection UUIDs) lives in
config/pipeline.yamlandconfig/mappings/*.yaml. To add a new document type, updatedoctypes.yaml; to update a collection UUID after a DSpace migration, updatecollections.yaml— no Python changes required. - Source priority for deduplication merging is defined in
config/pipeline.yaml → source_order. - The stylesheet lives in
ui/styles.css(pure CSS);app.pyonly injects colour tokens as CSS custom properties (var(--canard), etc.) via a small inline<style>block.
Each run produces a timestamped subfolder under data/:
data/
└── 2025-10-08_03-15/
├── Raw_WosItems.csv
├── Raw_ScopusItems.csv
├── Raw_CrossrefItems.csv
├── Raw_OpenalexItems.csv
├── Raw_ZenodoItems.csv
├── Raw_EpoItems.csv
├── DeduplicatedItems.csv
├── UnloadedItems.csv # clear duplicates found in Infoscience (discarded)
├── Items.csv
├── AuthorsAndAffiliations.csv
├── EpflAuthors.csv
├── ItemsWithOAMetadata.csv
├── ImportedItems.csv
├── RejectedItems.csv
└── Report_2025-10-08_03-15.xlsx
Run history, per-source statistics, publications, EPFL authors, and unit links are also written to DuckDB (data/pipeline_{env}.duckdb) and are browsable through the Streamlit UI.
The pipeline is intentionally stateless. Re-running it is always safe because:
- Sliding window — only publications within the configured date range are harvested.
- DSpace-aware dedup — the deduplicator queries Infoscience for existing items before loading, so already-imported records are never duplicated. Deduplication is type-aware:
- A dataset is only deduplicated against other datasets (scoped to the Datasets and Code collection); a title match in another collection is flagged but not discarded.
- A preprint whose published version already exists in Infoscience is forwarded to the DSpace workspace rather than silently dropped, so it can be reviewed and linked.
- Clear duplicates (same DOI, same type, same title + year within the same collection) are discarded without flagging.
- Stable
row_id— each record gets a deterministic hash of its key fields, ensuring consistent matching across runs.
Running daily with a 15-day window (the default) catches late-indexed publications while the overlap with previous windows is handled entirely by deduplication.
This project is licensed under the MIT License — see LICENSE for the full text.
The bundled DSpace REST Python Client (dspace/dspace_rest_client/) is licensed under the BSD 3-Clause License (© The Library Code GmbH). MIT and BSD-3-Clause are compatible permissive licences; the BSD-3-Clause copyright notice is preserved in the LICENSE file and in the bundled source.
If you use this software in your research or institutional work, please cite it using the metadata in CITATION.cff or the reference below:
@software{infoscience_import_pipeline,
author = {Sicot, Julien and Borel, Alain and Geoffroy, Géraldine},
title = {Infoscience Import Pipeline},
year = {2026},
publisher = {EPFL Library},
url = {https://github.com/epfllibrary/infoscience-imports},
license = {MIT}
}