An autonomous data curation agent for Synapse-backed data portals. Configurable for any disease domain or research focus via config/settings.yaml and config/keywords.yaml.
NADIA runs daily, discovers publicly available research datasets from scientific repositories, and provisions Synapse "pointer" projects for data manager review. It is powered by Claude Code and uses the Anthropic API for relevance scoring and annotation normalization.
The NF Data Portal configuration (neurofibromatosis / schwannomatosis) is the reference deployment. Set your own search terms, portal table IDs, Synapse team, and schema prefix in the config files to target a different disease domain.
flowchart TD
A([NADIA\nDaily trigger\nGitHub Actions]) --> B
subgraph DISCOVER["1 · Discovery"]
B[PubMed search\nMeSH + keywords\nlast 30 days] --> C[NCBI elink\npubmed → gds / sra / gap]
B --> D[Europe PMC annotations\nfull-text accession mining]
C --> E[Publication groups\nkeyed by PMID]
D --> E
E --> F[Secondary pass\nZenodo · Figshare · OSF\nPRIDE · ArrayExpress · PDC\nskip if PMID already found]
end
subgraph DEDUP["2 · Deduplication"]
E --> G{Match against\nNF Portal + agent\nstate table}
G -->|SKIP\nall accessions known| Z1([Log rejected_duplicate])
G -->|ADD\nnew accession for\nexisting study| H[Add Dataset to\nexisting project]
G -->|NEW\nno match| I[Create new project]
end
subgraph SCORE["3 · Score & Normalize"]
I --> J[Claude reasoning\nrelevance score 0–1\ndisease focus · assay · species]
J -->|score < 0.70\nor review / meta-analysis| Z2([Log rejected_relevance])
J -->|score ≥ 0.70| K[Fetch schema enums\nSynapse REST API]
K --> L[Claude reasoning\nnormalize raw metadata\n→ valid schema enum values]
end
subgraph BUILD["4 · Build Synapse Project"]
L --> M[Create Synapse project\nnamed after publication title]
M --> N[Create folders\nRaw Data · Source Metadata\n+ Processed Data]
N --> O[Per accession:\nfiles folder inside Raw Data\nor Processed Data]
O --> P[Enumerate files\nGEO FTP · ENA · Zenodo · PRIDE etc.]
P --> Q[File entities\npath= direct URL\nsynapseStore=False]
Q --> R[Per-file annotations\nassay · species · diagnosis\nspecimenID · fileFormat · …]
R --> S[Dataset entity\ndirect child of project\ncolumnIds = annotation fields]
S --> T[Link files as\nDataset items]
T --> U[Bind JSON Schema\nconfigurable URI prefix + template]
U --> V[Wiki page\nplain-language summary\n+ abstract + datasets table]
V --> AA[Self-audit\nverify all annotations · fix gaps\nvalidate schema compliance]
end
subgraph STATE["5 · State & Notify"]
AA --> W[Update state table\nProcessedStudies + RunLog\nstatus = synapse_created]
H --> W
W --> X[Create JIRA ticket\npending data manager review]
end
style DISCOVER fill:#1565c0,color:#fff,stroke:#1565c0
style DEDUP fill:#e65100,color:#fff,stroke:#e65100
style SCORE fill:#6a1b9a,color:#fff,stroke:#6a1b9a
style BUILD fill:#1b5e20,color:#fff,stroke:#1b5e20
style STATE fill:#880e4f,color:#fff,stroke:#880e4f
Each discovered publication becomes one Synapse project:
{Publication Title}/ ← Synapse Project
├── GEO_{AccessionID} Raw ← Dataset entity (Datasets tab, FASTQs)
├── GEO_{AccessionID} Processed ← Dataset entity (Datasets tab, processed files)
├── Zenodo_{AccessionID} ← Dataset entity (Datasets tab)
├── Raw Data/
│ └── GEO_{AccessionID}_files/ ← Folder — FASTQ File entities with direct URLs
├── Processed Data/
│ └── GEO_{AccessionID}_files/ ← Folder — processed files (e.g. Cell Ranger output)
└── Source Metadata/
- Dataset entities are direct children of the project so they appear in the portal's Datasets tab.
- File entities use
synapseStore=Falsewithpath=<direct download URL>— data is never uploaded to Synapse. - Each Dataset entity has annotation columns defined and is bound to the appropriate JSON schema (configured via
synapse.schema.uri_prefixandsynapse.schema.metadata_dictionary_urlinconfig/settings.yaml; the NF deployment uses the NF metadata dictionary).
├── CLAUDE.md Agent instructions (loaded by Claude Code at runtime)
├── lib/
│ ├── synapse_login.py Synapse authentication helper
│ ├── state_bootstrap.py Creates/retrieves agent state tables in Synapse
│ ├── keywords.yaml Disease domain search terms and PubMed MeSH query
│ ├── nf_keywords.yaml Legacy NF/SWN search terms (superseded by keywords.yaml)
│ └── settings.yaml Runtime configuration
├── prompts/
│ └── daily_task_template.md Task prompt for scheduled runs
├── config/ Environment-specific configuration
└── tests/ Unit tests
Note: Generated scripts are written to the
agent.workspace_dirpath (default:/tmp/nf_agent/) at runtime and are not committed to this repository.
- Python 3.12+
- A Synapse account with write access to your agent state project
- An Anthropic API key (Claude Code / claude CLI)
- NCBI API key (optional, increases rate limit from 3 → 10 req/s)
pip install -r lib/requirements.txt| Variable | Required | Purpose |
|---|---|---|
SYNAPSE_AUTH_TOKEN |
Yes | Synapse personal access token |
ANTHROPIC_API_KEY |
Yes | Anthropic API key for Claude |
STATE_PROJECT_ID |
Yes | Synapse project ID for agent state tables |
NCBI_API_KEY |
Recommended | NCBI Entrez API key |
JIRA_BASE_URL |
Optional | e.g. https://sagebionetworks.jira.com |
JIRA_USER_EMAIL |
Optional | JIRA service account email |
JIRA_API_TOKEN |
Optional | JIRA API token |
Create a Synapse project to hold the agent's state tables. The agent will auto-create two tables on first run, named using the agent.state_table_prefix from config/settings.yaml (default: NF_DataContributor):
{prefix}_ProcessedStudies— tracks every accession processed{prefix}_RunLog— one row per daily run
Set STATE_PROJECT_ID to the syn ID of that project.
export SYNAPSE_AUTH_TOKEN=...
export ANTHROPIC_API_KEY=...
export STATE_PROJECT_ID=syn...
export NCBI_API_KEY=... # optional
claude --permission-mode bypassPermissions \
--add-dir /path/to/nf-data-contributor \
-p "$(cat prompts/daily_task_template.md)"See GitHub setup instructions below.
-
Fork or clone this repository into your GitHub organization.
-
Add repository secrets (Settings → Secrets and variables → Actions → New repository secret):
Secret name Value SYNAPSE_AUTH_TOKENSynapse personal access token for the bot account ANTHROPIC_API_KEYAnthropic API key STATE_PROJECT_IDSynapse project ID for state tables (e.g. syn74273218)NCBI_API_KEYNCBI API key (recommended) JIRA_BASE_URLOptional — JIRA base URL JIRA_USER_EMAILOptional — JIRA service account email JIRA_API_TOKENOptional — JIRA API token -
Create the workflow file at
.github/workflows/daily_run.yml:name: NADIA — Daily Run on: schedule: - cron: '0 8 * * *' # 08:00 UTC daily workflow_dispatch: # allow manual trigger jobs: run-agent: runs-on: ubuntu-latest timeout-minutes: 120 steps: - uses: actions/checkout@v4 - name: Set up Python uses: actions/setup-python@v5 with: python-version: '3.12' - name: Install dependencies run: pip install -r lib/requirements.txt - name: Install Claude Code CLI run: npm install -g @anthropic-ai/claude-code - name: Run agent env: SYNAPSE_AUTH_TOKEN: ${{ secrets.SYNAPSE_AUTH_TOKEN }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} STATE_PROJECT_ID: ${{ secrets.STATE_PROJECT_ID }} NCBI_API_KEY: ${{ secrets.NCBI_API_KEY }} JIRA_BASE_URL: ${{ secrets.JIRA_BASE_URL }} JIRA_USER_EMAIL: ${{ secrets.JIRA_USER_EMAIL }} JIRA_API_TOKEN: ${{ secrets.JIRA_API_TOKEN }} AGENT_REPO_ROOT: ${{ github.workspace }} run: | claude --permission-mode bypassPermissions \ --max-turns 120 \ --output-format text \ --add-dir ${{ github.workspace }} \ -p "$(cat prompts/daily_task_template.md)"
-
Enable Actions in your repository (Settings → Actions → Allow all actions).
-
Test with a manual trigger: Go to Actions → NADIA — Daily Run → Run workflow.
The agent operates under strict safety rules defined in CLAUDE.md:
- The three NF Data Portal tables (
syn52694652,syn16858331,syn16859580) are read-only — the agent never mutates portal data. - The agent only writes to projects it created in the current run and to its own state tables.
- Maximum 50 Synapse write operations per run.
- All created projects have
resourceStatus=pendingReview— a human data manager must approve before they appear publicly on the portal.