A curated list of biological databases with AI Agent-friendliness ratings. Built for AI for Science researchers and developers. Contributions welcome!
- π·οΈ Legend
- 𧬠Sequence & Genome
- π§ͺ Protein & Structure
- πΏ Plants
- π¦ Microbes & Viruses
- π¬ Single Cell
- βοΈ Enzymes & Metabolism
- π Compounds & Drugs
- π₯ Disease & Clinical
- π Interaction & Network
- π Expression & Regulation
- π€ Contributing
- π‘ API Reference
- π License
| Icon | Meaning | Criteria |
|---|---|---|
| π€ | Agent-Ready | Public REST/GraphQL API + bulk data download available |
| Partial | Either API or download available, but not both | |
| β | Manual | Web UI only; no API and no bulk download |
| π | API Key Required | Free registration needed for API access (still Agent-Ready) |
| π² | Commercial | Paid API or data access required (downgraded to Partial) |
| Name | Description | Agent | Links |
|---|---|---|---|
| 𧬠DDBJ | DNA Data Bank of Japan β primary nucleotide sequence archive and INSDC member alongside GenBank and ENA, providing annotated sequence data from Japanese research institutions. Covers all published nucleotide sequences with metadata including taxonomy, source organisms, and functional annotations, synchronized daily across the International Nucleotide Sequence Database Collaboration. | π€ | REST API, FTP |
| 𧬠ENA | European Nucleotide Archive β comprehensive nucleotide sequence repository at EMBL-EBI and INSDC member covering all domains of life. Provides assembled genomes, raw reads, transcriptome assemblies, and functional annotation for >200 TB of sequence data with the ENA Browser for sequence lookup and CRAM/FASTQ download via REST API. | π€ | REST API, FTP |
| 𧬠ENCODE | Encyclopedia of DNA Elements β comprehensive catalog of 13,000+ functional genomics experiments including ChIP-seq, RNA-seq, ATAC-seq, DNase-seq, and Hi-C across human and mouse cell lines and tissues. Provides the Registry of candidate cis-Regulatory Elements (cCREs) with integrated chromatin accessibility, histone modification, and transcription factor binding data. Essential reference for genome regulatory element annotation and deep learning-based sequence-to-function models. | π€ | REST API, Downloads |
| 𧬠Ensembl | Genome browser providing comprehensive gene annotation for 300+ vertebrate and non-vertebrate species with automated and manual curation pipelines. Integrates variation data (SNPs, indels, structural variants), comparative genomics (gene trees, whole-genome alignments, synteny), and regulatory annotation across all species. The REST API supports programmatic access to variant effect prediction (VEP), sequence retrieval, and homology queries. | π€ | REST API, FTP |
| 𧬠GenBank | NIH genetic sequence database and founding INSDC member β contains >2 billion annotated DNA sequences from >400,000 organisms with daily submissions from individual laboratories and large-scale sequencing projects. Provides the authoritative GenBank flatfile format with locus, definition, accession, features (CDS, exons, regulatory regions), and reference to published literature. Updated every two months (release cycles) with newly submitted and revised entries. | π€ | E-Utilities API, FTP |
| 𧬠GENCODE | High-quality reference gene annotation for human (GRCh38) and mouse (GRCm39) genomes produced by the ENCODE consortium. Integrates manual annotation from the HAVANA team with automated Ensembl annotation pipelines, providing consistent transcript models, pseudogenes, long non-coding RNAs, and small non-coding loci. Widely used as the authoritative gene set in ENCODE data analysis pipelines and genome interpretation tools. | FTP downloads, no API | |
| 𧬠gnomAD | Genome Aggregation Database (gnomAD v4) β aggregated exome and whole-genome sequencing data from 141,456 individuals including 76,156 genomes across diverse populations. Provides allele frequencies, constraint metrics (pLI, LOEUF), population-specific variant quality annotations, and structural variant calls processed through a standardized pipeline. The gold-standard reference for population variant frequency filtering in rare disease diagnostics and human genetics research. | π€ | GraphQL API, Downloads |
| 𧬠HGNC | HUGO Gene Nomenclature Committee β the sole authority for assigning standardized gene symbols and names to human protein-coding genes, non-coding RNA genes, and pseudogenes. Maintains >42,000 approved gene symbols with cross-references to all major databases (NCBI, Ensembl, UniProt, OMIM) and provides gene family classifications, selectable markers, and locus-specific database links. Essential for unambiguous gene identification in genomic analysis and cross-database data integration. | π€ | REST API, Downloads |
| 𧬠NCBI Assembly | NCBI Assembly database providing access to >1 million assembled genome sequences with standardized RefSeq accession identifiers, assembly statistics (contig N50, total length, gap count), and genome coverage metadata. Supports programmatic retrieval via E-Utilities with filters by organism, assembly level (complete genome, scaffold, contig), and submitter. The primary reference for genome assembly naming and versioning across all NCBI resources. | π€ | E-Utilities API, FTP |
| 𧬠NCBI Gene | NCBI Gene database providing comprehensive gene-centric records for >10 million species with integrated access to RefSeq sequences, genomic context, expression data, conserved domains, and protein information. Links every annotated gene to PubMed literature, variation resources (dbSNP, ClinVar), and pathway databases with stable GeneID identifiers. Core infrastructure for connecting gene symbols across NCBI's interconnected data ecosystem. | π€ | E-Utilities API, FTP |
| 𧬠PharmGKB | Pharmacogenomics Knowledge Base β expertly curated database of how genetic variation affects drug response phenotypes, covering >4,000 clinically actionable variant-drug associations. Includes clinical guideline annotations from CPIC and FDA-approved drug labels with pharmacogenomics information, dosing recommendations, and curated pharmacokinetic pathway diagrams. The primary reference for building pharmacogenomic AI models and clinical decision support systems. | π€ | REST API, Downloads |
| 𧬠RefSeq | NCBI Reference Sequence Database β the gold-standard curated, non-redundant collection of genomic DNA, transcript (mRNA), and protein sequences for >90,000 organisms. RefSeq records are manually reviewed by NCBI staff with consistent accession numbering (NC_, NM_, NP_ prefixes) and stable version tracking for reproducible scientific analysis. Essential reference for RNA-seq quantification, variant annotation, and comparative genomics benchmarks. | π€ | E-Utilities API, FTP |
| 𧬠SRA | NCBI Sequence Read Archive β the largest public repository of raw high-throughput sequencing data with >40 petabases of sequence from Illumina, PacBio, Oxford Nanopore, and other platforms. Accepts data from any sequencing study type (whole-genome, RNA-seq, ChIP-seq, metagenomic, single-cell) under the INSDC umbrella with standardized metadata. Primary source for training deep learning models on raw sequencing signals and reanalyzing published datasets. | π€ | E-Utilities API, FTP |
| 𧬠T2T Genome | Telomere-to-Telomere consortium β first complete, gapless assembly of a human genome (CHM13, published in Science 2022), adding ~200 Mbp of previously unresolved sequence including centromeric satellites, segmental duplications, and ribosomal DNA arrays. Provides the first complete annotation of all five acrocentric short arms, enabling studies of centromere biology and previously inaccessible repetitive genomic regions. Represents the new benchmark reference for human genome analysis, surpassing GRCh38 in completeness. | Downloads, no API | |
| 𧬠UCSC Genome Browser | Interactive genome browser hosting 100+ annotation tracks for human and 100+ model organism genomes with deep integration of ENCODE, Roadmap Epigenomics, and comparative genomics data. Provides BLAT alignment for rapid sequence similarity search, the Table Browser for programmatic track data access, and a REST API for custom track upload and data retrieval. The most widely used platform for visual genome exploration and educational tool for genome annotation. | REST API, FTP downloads | |
| 𧬠UniProt | UniProt Knowledgebase (UniProtKB) β the world's most comprehensive protein sequence and functional information resource, comprising manually reviewed Swiss-Prot (>570,000 entries) and automatically annotated TrEMBL (>250 million entries). Covers protein function, domain architecture, subcellular localization, post-translational modifications, variant effects, and cross-references to 170+ external databases. The foundational protein reference for bioinformatics pipelines, from sequence search to functional enrichment and AI-based protein structure prediction. | π€ | REST API, FTP |
| Name | Description | Agent | Links |
|---|---|---|---|
| π§ͺ AlphaFold DB | Predicted protein structures for 200M+ proteins from AlphaFold2, hosted by EMBL-EBI, covering most entries in UniProt across all domains of life. Provides per-residue confidence scores (pLDDT) and predicted aligned error (PAE) metrics for model quality assessment, enabling reliable structure-based function prediction. The single most transformative AI-generated structural biology resource, widely used for protein design, docking, and functional annotation at proteome scale. | π€ | API docs, Bulk download |
| π§ͺ CATH | Hierarchical classification of protein domain structures into Class, Architecture, Topology, and Homologous superfamily levels (CATH v4.3, 150M+ domain assignments). Provides sequence and structure search tools (CATH-SCOP) with functional family annotations (FunFams) for evolutionary and functional inference. Used for protein structure-function studies, domain boundary prediction, and evolutionary analysis of protein families. | π€ | REST API, Downloads |
| π§ͺ ELM | Eukaryotic Linear Motif resource β curated database of experimentally validated short linear motifs (SLiMs) in eukaryotic proteins, typically 3-10 residues mediating transient protein-protein interactions. Provides motif consensus patterns, structural context, instance annotations, and cross-references to PDB structures for motif-mediated interaction interfaces. Key resource for studying signaling networks, post-translational modification sites, and cell compartment targeting sequences. | Downloads, web API limited | |
| π§ͺ Human Protein Atlas | Protein expression atlas across 44 human tissues, 20 cancer types, and 47 cell lines combining antibody-based immunohistochemistry (Tissue Atlas) with transcriptomics (RNA-seq) and mass spectrometry-based proteomics. Organized into Tissue Atlas, Cell Atlas (subcellular protein localization), and Pathology Atlas (cancer survival associations) sub-projects with >26,000 antibodies targeting the human proteome. Essential resource for tissue-specific protein expression analysis and biomarker discovery. | π€ | REST API, Downloads |
| π§ͺ InterPro | Integrative protein domain and family classification system combining 13 member databases (Pfam, PROSITE, SMART, PRINTS, CDD, and others) into a unified annotation framework. Provides protein sequence analysis with functional annotations, domain architectures, and GO term associations for >2 million sequences across all kingdoms of life. The authoritative resource for automated protein functional classification in genome annotation pipelines. | π€ | REST API, FTP |
| π§ͺ MobiDB | Database of intrinsically disordered proteins (IDPs) and regions (IDRs) with consensus disorder predictions from 10+ computational predictors including IUPred, ESpritz, and DeepCNF. Provides per-residue disorder scores, predicted secondary structure, protein-binding regions (MoRFs), and functional annotations linked to disordered regions. Essential resource for studying phase separation, unstructured protein function, and disorder-based drug targeting. | π€ | REST API, Downloads |
| π§ͺ PDB | Worldwide Protein Data Bank (wwPDB) β the global repository of >220,000 experimentally determined 3D structures of proteins, nucleic acids, and complexes, solved by X-ray crystallography, cryo-electron microscopy, and NMR spectroscopy. Provides standardized mmCIF format with full atomic coordinates, experimental metadata, structure factors, and validated biological assembly information. The gold-standard resource for structural biology, structure-based drug design, and AI training for protein structure prediction. | π€ | REST API, FTP download |
| π§ͺ Pfam | Protein families database (now part of InterPro) providing >19,000 curated multiple sequence alignments and hidden Markov model (HMM) profiles for protein domain families across all domains of life. Each family includes seed alignments, full alignments, consensus sequences, and clan-level groupings of evolutionarily related families. Widely used for domain annotation of novel proteins, metagenomic functional profiling, and deep learning feature encoding in protein property prediction. | π€ | InterPro API, FTP |
| π§ͺ PROSITE | Database of protein domains, families, and functional sites using manually curated patterns (regular expressions) and profiles (position-specific scoring matrices) to identify conserved sequence regions. Covers >1,300 documented entries with annotations for active sites, binding sites, post-translational modification sites, and signature motifs associated with specific protein families. Useful for functional annotation of newly sequenced proteins and exploring protein family conservation. | Downloads via FTP, no public API | |
| π§ͺ SAbDab | Structural Antibody Database β all publicly available antibody and nanobody structures from the PDB with consistent annotation of complementarity-determining regions (CDRs), framework regions, antigen partners, and experimental method. Provides sequence-level clustering, structural statistics (VL/VH packing angles, CDR loop conformations), and therapeutic antibody design tools including SAbPred. The primary resource for antibody structure analysis, engineering, and AI-based antibody design. | π€ | REST API, Downloads |
| π§ͺ SCOP | Structural Classification of Proteins β manually curated hierarchical classification of protein domains based on evolutionary relationships and structural similarity, organized by Class, Fold, Superfamily, and Family levels. Provides structural annotations for >140,000 PDB entries with domain boundaries, fold descriptions, and species-specific counts of domain occurrences. Key resource for studying protein fold evolution, domain architecture analysis, and structural genomics targets. | Downloads, web API limited | |
| π§ͺ SWISS-MODEL Repository | Homology-modelled protein structures covering >2 million models for model organism proteomes using SWISS-MODEL's automated comparative modeling pipeline. Provides per-residue model quality scores (QMEAN, GMQE), oligomeric state annotations, and ligand information derived from template PDB structures. Essential resource for structural genomics, protein function prediction, and structure-based variant effect analysis when experimental structures are unavailable. | π€ | REST API, Downloads |
| Name | Description | Agent | Links |
|---|---|---|---|
| πΏ BAR | Bio-Analytic Resource for Plant Biology β curated collection of web-based tools and datasets for Arabidopsis thaliana and crop species, including the Arabidopsis eFP Browser (gene expression visualization across tissues, developmental stages, and stress conditions) and the Arabidopsis Interactions Viewer. Provides expression data from >10,000 Affymetrix ATH1 arrays and RNA-seq experiments with clustering and co-expression analysis pipelines. Key resource for plant functional genomics and gene discovery. | Downloads, mostly web-based | |
| πΏ Ensembl Plants | Ensembl genome browser specialized for plant genomes, covering 67+ species including major crops (rice, maize, wheat, soybean, tomato, grape) and model plants (Arabidopsis, Brachypodium). Provides gene annotation, variation data (SNPs, structural variants), comparative genomics (gene trees, whole-genome alignments), and regulatory features using the Ensembl annotation pipeline. Shares the same REST API and VEP (Variant Effect Predictor) infrastructure as main Ensembl, enabling cross-species comparative genomics. | π€ | REST API, FTP |
| πΏ Gramene | Comparative plant genomics and pathway resource for crop plants, powered by Ensembl Plants infrastructure for genome browsing and annotation. Integrates genetic variation data (SNPs, QTLs, phenotype associations), metabolic pathways from Plant Reactome, and gene ontology annotations across rice, maize, wheat, sorghum, and other cereals. Provides a comprehensive orthology-based framework for translating functional genomics findings between model plants and crops. | π€ | REST API, Downloads |
| πΏ MaizeGDB | Comprehensive maize (Zea mays) genetics and genomics database serving the maize research community with the B73 reference genome, annotations, and genetic variation data. Provides curated phenotypic data for mutant stocks, quantitative trait loci (QTLs), genetic maps, and expression profiles integrated with the genome browser. The central community resource for maize as a C4 photosynthesis model and major crop species. | Downloads, API limited | |
| πΏ Phytozome | JGI/DOE plant comparative genomics portal providing assembled genomes and annotations for 100+ plant species spanning algae, bryophytes, ferns, gymnosperms, and angiosperms. Features curated gene family trees, whole-genome duplications, synteny blocks, and functional annotations (KEGG, GO, Pfam domains) for every genome. The primary resource for DOE-relevant plant and algal genomics, bioenergy crop research, and plant evolutionary biology. | Bulk download, no public REST API | |
| πΏ Plant Reactome | Curated plant metabolic and signaling pathways using the Reactome framework, covering rice, maize, wheat, soybean, and Arabidopsis with >800 pathway diagrams. Provides orthology-based pathway projections from human and model organism reactions, pathway enrichment analysis tools, and cross-references to Gramene and Ensembl Plants gene annotations. Enables comparative plant metabolism studies and pathway-based interpretation of plant multi-omics data. | π€ | ContentService API, Downloads |
| πΏ PlantGDB | Plant Genome Database providing expressed sequence tag (EST) assemblies, spliced alignments, and genome browser views for >100 diverse plant species. Includes resources for gene structure prediction, transcript assembly, and alternative splicing annotation based on EST and cDNA sequence evidence. Useful for cross-species comparative transcriptomics and gene model validation in non-model plant species. | FTP downloads, no API | |
| πΏ PlantTFDB | Plant Transcription Factor Database β comprehensive collection of >320,000 transcription factors across 165 plant species, classified into 58 families based on conserved DNA-binding domains. Provides domain architecture analysis, phylogenetic trees, expression profiles, and cross-species orthology assignments for each TF family. Primary resource for plant regulatory genomics, TF family evolutionary studies, and building plant gene regulatory network models. | Downloads, web-based only | |
| πΏ PLAZA | Plant comparative genomics platform providing gene family classification, orthology inference, and synteny analysis across 37+ sequenced plant genomes (PLAZA 5.0). Features precomputed gene families with multiple sequence alignments and phylogenetic trees, whole-genome synteny blocks, and integrated functional annotations (GO, InterPro, KEGG). Enables evolutionary analyses of gene family expansion, duplication events, and conserved non-coding sequences across plants. | Bulk downloads, no API | |
| πΏ PlncRNADB | Plant long non-coding RNA database cataloging lncRNAs from >40 plant species with expression profiles, epigenetic marks (histone modifications, DNA methylation), and functional annotations. Provides sequence information, genomic coordinates, tissue-specific expression patterns, and predicted secondary structures for each lncRNA transcript. Useful for studying plant lncRNA biology, regulatory interactions, and epigenetic regulatory mechanisms. | β | Web-based with search |
| πΏ RAP-DB | Rice Annotation Project Database β high-quality annotation of the Oryza sativa japonica (Nipponbare) reference genome by the International Rice Genome Sequencing Project (IRGSP). Provides manually curated gene models, transcript variants, functional annotations (GO, InterPro), expression data, mutant lines, and miRNA/target predictions. The authoritative rice genome annotation resource for the premier monocot model organism and staple crop feeding billions. | Downloads, no API | |
| πΏ SoyBase | Soybean genetics and genomics database for Glycine max, providing the Williams 82 reference genome with gene annotations, genetic maps, and marker data. Curates QTLs for agronomic traits (yield, disease resistance, seed composition), mutant phenotypes, and breeding germplasm information with integrated genome browser and BLAST search. Primary community resource for soybean researchers and legume comparative genomics. | Downloads, no public API | |
| πΏ TAIR | The Arabidopsis Information Resource β comprehensive model plant database for Arabidopsis thaliana, the dicot model organism, providing annotation for >27,000 protein-coding genes with structured descriptions and gene ontology (GO) assignments. Curates gene function data from >50,000 publications including expression patterns, mutant phenotypes, protein interactions, and subcellular localizations. The foundational reference for plant functional genomics and the primary model for understanding plant gene function. | REST API, bulk download via request | |
| πΏ WheatOmics | Wheat multi-omics database integrating transcriptome, proteome, metabolome, and epigenome data for common wheat (Triticum aestivum). Provides gene expression profiles across tissues, developmental stages, and stress conditions, along with protein abundance, metabolite profiles, and DNA methylation data. Enables integrative analysis of the large hexaploid wheat genome for understanding complex trait biology and stress responses. | β | Web-based query only |
| Name | Description | Agent | Links |
|---|---|---|---|
| π¦ BacDive | Bacterial Diversity database β DSMZ-curated resource providing standardized metadata for 90,000+ bacterial and archaeal strains with >600 metadata fields per strain covering morphology, physiology, cultivation conditions, metabolism (oxygen, enzymes, carbon sources), antibiotic sensitivity, and isolation source. Cross-references each strain to 16S rRNA sequences, genome assemblies, and taxonomic classifications. Essential resource for microbiological strain characterization and comparative microbial phenomics. | π€ | REST API, Downloads π |
| π¦ BV-BRC | Bacterial and Viral Bioinformatics Resource Center β NIAID-funded bioinformatics resource integrating genomic and metadata for >600,000 bacterial and >9 million viral pathogen strains. Provides comprehensive analysis pipelines including genome assembly, annotation, phylogenetic tree building, and comparative genomics with a unified REST API. Primary resource for infectious disease research, pathogen surveillance, and antimicrobial resistance tracking. | π€ | REST API, Downloads |
| π¦ CARD | Comprehensive Antibiotic Resistance Database β expertly curated collection of antibiotic resistance genes (ARGs), their molecular mechanisms, and associated antibiotics, organized through the Antibiotic Resistance Ontology (ARO). Includes the Resistance Gene Identifier (RGI) tool for detecting ARGs in genomic and metagenomic data using BLAST and HMM-based models with resistance mechanism prediction. The standard reference for antimicrobial resistance gene annotation and resistome analysis in pathogen and microbiome studies. | π€ | REST API, Downloads |
| π¦ GISAID | Global Initiative on Sharing All Influenza Data β the primary global repository for influenza and SARS-CoV-2 genome sequences, hosting >15 million SARS-CoV-2 sequences and extensive influenza data. Enforces a unique data sharing agreement requiring attribution and collaboration, which has driven unprecedented global pathogen sequence sharing during the COVID-19 pandemic. Indispensable for pandemic surveillance, variant tracking (e.g., WHO Pango lineages), and viral evolution monitoring. | π API requires registration and agreement; Data access | |
| π¦ Greengenes2 | Updated 16S rRNA reference database with whole-genome-enhanced taxonomic backbone integrating >50,000 genomes from the Web of Life and GTDB taxonomies. Provides full-length and region-specific 16S sequences with consistent phylogenetic placement, replacing the earlier Greengenes with genome-scale taxonomic resolution. The preferred 16S reference for QIIME 2-based microbiome analysis with improved classification accuracy. | π€ | Downloads + API via Qiita/FTP |
| π¦ IMG/M | Integrated Microbial Genomes & Microbiomes β JGI/DOE metagenome analysis platform providing functional annotation for >50,000 metagenomes and >100,000 microbial genomes with the JGI annotation pipeline (IMG-Pipeline). Features genome-centric analysis (phylogenetic distribution, gene clusters, biosynthetic gene clusters) and metagenome-centric analysis (binning, assembly, functional profiling) with cross-metagenome comparison tools. Premier resource for large-scale comparative metagenomics and microbial functional genomics. | REST API, data download via JGI portal | |
| π¦ MGnify | EBI metagenomics resource providing free, standardized analysis and archiving of microbiome-derived sequences from >200,000 samples (amplicon, metagenomic, metatranscriptomic). Offers automated pipelines for taxonomic profiling (SILVA, ITS) and functional annotation (InterPro, Pfam, KEGG) with public access to all analysis results. Key resource for large-scale microbiome data reanalysis and meta-analysis studies. | π€ | REST API, Downloads |
| π¦ MicrobiomeDB | Multi-omics microbiome data warehouse integrating 16S rRNA, metagenomic, metatranscriptomic, and metabolomic data from >100 published microbiome studies across human body sites and environmental contexts. Provides interactive analysis tools including differential abundance testing, alpha/beta diversity analysis, and correlation networks with a unified data model. Enables comparative microbiome analysis and hypothesis generation across previously isolated study datasets. | Downloads, API limited | |
| π¦ SILVA | Comprehensive ribosomal RNA (rRNA) gene database containing >9 million aligned small subunit (SSU/16S/18S) and large subunit (LSU/23S/28S) sequences for bacteria, archaea, and eukarya. Provides curated, quality-checked alignments with curated taxonomic classifications using the SILVA ACT (Alignment, Classification, Tree) pipeline, updated semi-annually. The most widely used reference for rRNA-based taxonomic profiling, phylogenetic placement, and microbiome diversity analysis. | Downloads, no public API | |
| π¦ VFDB | Virulence Factor Database β comprehensive collection of experimentally validated bacterial virulence factors (VFs) from >30 clinically important pathogenic genera, organized by pathogen species and virulence factor category (adhesion, toxin, secretion system, biofilm). Provides core and accessory VF datasets, DNA and protein sequences, and BLAST search for virulence gene identification in newly sequenced pathogens. Primary resource for bacterial pathogenesis research and virulence determinant detection. | Downloads, web-based API | |
| π¦ ViPR | Virus Pathogen Resource β NIAID-funded bioinformatics resource integrating genomic, proteomic, and metadata for human pathogenic viruses including influenza, coronaviruses, flaviviruses, filoviruses, and paramyxoviruses. Provides genome assembly and annotation, 3D protein structure models, immune epitope mapping, and phylogenetic analysis tools (now merged into BV-BRC). Essential resource for viral genomics, vaccine design research, and outbreak response. | Downloads, API via BV-BRC |
| Name | Description | Agent | Links |
|---|---|---|---|
| π¬ Azimuth | HuBMAP-powered reference-based automated cell type annotation platform for single-cell RNA-seq data, providing precomputed references for human PBMC, bone marrow, brain, heart, kidney, lung, and other tissues. Uses a Seurat-based reference mapping pipeline (findTransferAnchors + MapQuery) to label query cells and project them onto a UMAP embedding. Enables rapid, standardized cell type annotation for new scRNA-seq datasets. | π€ | API via R/Python packages, Downloads |
| π¬ CZ CELLxGENE | Chan Zuckerberg Initiative single-cell data platform curating >1,000 scRNA-seq datasets with 50M+ cells across human and model organisms, all processed into standardized h5ad and Seurat formats with consistent cell type annotations. Provides the Census API for programmatic access to the entire corpus as a unified cell-by-gene matrix, enabling cross-dataset analysis and large-scale machine learning. The fastest growing single-cell data resource and the primary platform for AI-ready single-cell data. | π€ | REST API, Downloads |
| π¬ DISCO | Database of Immune Cells integrating scRNA-seq expression profiles, eQTL data, and cell-type-specific gene regulation for immune cell types. Provides interactive exploration of gene expression across immune cell subtypes, identification of cell-type-specific marker genes, and linkage of genetic variants to immune gene expression. Specialized resource for immunogenomics and understanding genetic regulation of immune cell states. | Downloads, web-based | |
| π¬ Human Cell Atlas | International consortium building comprehensive reference maps of all human cell types, with >30 million cells profiled across 18 biological networks (tissue systems) using scRNA-seq, scATAC-seq, and spatial transcriptomics. Provides a Data Portal with standardized metadata and matrix downloads for all contributed datasets, enabling cross-tissue and cross-assay integration. Foundational resource for defining the human cell type landscape and building AI models of human cellular identity. | π€ | Data Portal API, Downloads |
| π¬ JingleBells | Repository of standardized scRNA-seq datasets in BAM format specifically designed for benchmarking immune repertoire analysis tools. Provides raw sequencing data from 10x Genomics platforms for T-cell receptor (TCR) and B-cell receptor (BCR) repertoire studies, with standardized metadata and ground-truth annotations. Essential evaluation benchmark for developing and comparing immune repertoire analysis algorithms from single-cell data. | Downloads, no API | |
| π¬ PanglaoDB | Single-cell transcriptomics resource aggregating processed scRNA-seq data for mouse and human from public repositories, with precomputed cell type marker genes and expression matrices. Provides a gene expression search interface, cell type enrichment analysis, and downloadable count matrices with cell cluster annotations. Useful for identifying cell type markers and exploring gene expression patterns across published scRNA-seq studies. | Downloads, no API | |
| π¬ scRNASeqDB | Human single-cell RNA-seq database providing gene expression profiles across diverse conditions and cell types from published studies, with search by gene, cell type, or condition. Provides visualization of expression distributions and detection rates for queried genes across different cell clusters. Useful for rapidly assessing cell-type-specific gene expression patterns without reprocessing raw data. | β | Web-based query only |
| π¬ Single Cell Expression Atlas | EMBL-EBI single-cell expression atlas providing standardized reprocessing of scRNA-seq datasets across species (human, mouse, zebrafish, and others) using consistent pipelines (iSEE, clustering, marker identification). Provides baseline expression across cell types, differential expression between conditions, and cell-type-specific marker gene identification with a uniform analytical framework. Enables cross-study and cross-species comparison of single-cell expression data. | π€ | REST API, Downloads |
| π¬ Single Cell Portal | Broad Institute's single-cell data portal hosting curated scRNA-seq studies with interactive exploration, clustering visualization, differential expression analysis, and cell type annotation. Provides downloadable count matrices, metadata, and analysis results for each study with integration to the Broad's genomic data ecosystem (DepMap, GTEx). Key platform for sharing and reanalyzing high-quality single-cell datasets produced by Broad-affiliated consortia. | API, downloads per study | |
| π¬ Tabula Sapiens | CZ Biohub cell atlas profiling ~500,000 cells from 24 human organs across multiple donors using scRNA-seq, with >400 annotated cell types including rare and previously undescribed populations. Provides cell-type-specific gene expression, differentially expressed genes per tissue, and matched microbiome data from the same donors. Unique resource for cross-tissue human cell type comparison and studying cell-type-specific gene regulation across the entire body. | Downloads, limited API | |
| π¬ TISCH | Tumor microenvironment single-cell atlas integrating >2 million cells from 79 cancer types across 300+ published scRNA-seq datasets, with standardized cell type annotation (malignant, immune, stromal). Provides interactive gene expression heatmaps, cell-type composition analysis, differential expression between tumor vs. normal, and survival association analysis. Comprehensive resource for tumor immunology, cancer cell state discovery, and immune checkpoint target identification. | Downloads, API via gene search | |
| π¬ TISSUES | Tissue-disease-gene associations database integrating evidence from transcriptomics (RNA-seq, microarray), proteomics (mass spectrometry), and automated text mining of PubMed abstracts. Provides confidence scores for tissue-gene associations and tissue-specific expression patterns across human and major model organisms. Enables rapid query of where a gene is expressed and which tissues are associated with specific diseases. | Downloads, no API |
| Name | Description | Agent | Links |
|---|---|---|---|
| βοΈ BioCyc | Collection of 20,000+ Pathway/Genome Databases (PGDBs) for sequenced organisms, generated using the Pathway Tools software from annotated genomes. Each PGDB includes computationally predicted metabolic pathways and a subset (e.g., EcoCyc, HumanCyc) with manual curation from literature. Provides pathway enrichment analysis, metabolic maps, and regulatory network visualization for systems biology studies. | REST API, API key required; downloads limited | |
| βοΈ BRENDA | Comprehensive enzyme information system containing >40 million data points manually extracted from >140,000 scientific publications covering >8,400 EC classes. Provides detailed kinetic parameters (K_m, k_cat, V_max), enzyme structures, reaction mechanisms, substrates/products, inhibitors, cofactors, and organism-specific enzyme occurrence. The most comprehensive enzyme knowledge base globally, indispensable for enzyme engineering, metabolic modeling, and structure-function studies. | SOAP API, bulk download via academic license only | |
| βοΈ CAZy | Carbohydrate-Active enZYmes database β curated modular classification of enzymes involved in carbohydrate metabolism, organized into >300 families across Glycoside Hydrolases (GH), GlycosylTransferases (GT), Polysaccharide Lyases (PL), Carbohydrate Esterases (CE), Auxiliary Activities (AA), and Carbohydrate-Binding Modules (CBM). Tracks sequence-structure-function relationships for carbohydrate-active enzymes with family-specific HMM profiles. Essential resource for biofuel research, gut microbiome enzyme discovery, and plant cell wall degradation studies. | Downloads via FTP request, no public API | |
| βοΈ eQuilibrator | Biochemical thermodynamics calculator providing standardized Gibbs free energy estimates for biochemical reactions under physiological conditions (pH, ionic strength, temperature). Covers >70,000 compounds and >80,000 reactions with group contribution method predictions and experimentally-derived thermodynamic parameters. Essential for metabolic engineering, pathway feasibility analysis, and synthetic biology design. | π€ | REST API, Downloads |
| βοΈ ExPASy ENZYME | Swiss-Prot enzyme nomenclature database providing the authoritative IUBMB Enzyme Nomenclature (EC number system) with accepted names, substrates, products, cofactors, and cross-references to UniProt entries for every classified enzyme. Links each EC class to known protein sequences and provides information on enzyme regulation, tissue distribution, and disease associations. The official reference for enzyme classification in bioinformatics and biochemistry. | Downloads via FTP, no API | |
| βοΈ HMDB | Human Metabolome Database β comprehensive catalog of >220,000 human metabolite entries with detailed chemical, clinical, and biochemical information. Each entry includes NMR/MS spectra, concentration ranges in biofluids, disease associations, metabolic pathway links (KEGG, MetaCyc), and enzyme/protein targets. The primary reference for human metabolomics, biomarker discovery, and integrating metabolomics data with systems biology models. | Downloads, no public API | |
| βοΈ IntEnz | Integrated relational Enzyme database at EMBL-EBI providing the official IUBMB enzyme nomenclature with accepted reactions, substrates, products, cofactors, and cross-references to ChEBI for chemical participants. Covers all EC classes with structured XML-formatted data including literature citations for each enzyme entry. Authoritative resource for resolving EC number hierarchy, enzyme-reaction mapping, and biochemical ontology references. | FTP downloads, SOAP API deprecated | |
| βοΈ KEGG | Kyoto Encyclopedia of Genes and Genomes β integrated systems biology resource covering >500 reference metabolic and signaling pathways from >40 reference organisms with hierarchical KEGG MODULE, KEGG ORTHOLOGY (KO), and BRITE functional classifications. Integrates pathway maps, drug information, disease annotations, and genomic data into KO-based cross-species functional prediction. Foundational resource for pathway enrichment analysis, metabolic network reconstruction, and functional interpretation of omics data. | π€ | REST API, FTP π² |
| βοΈ Lipid Maps | LIPID MAPS Structure Database β the largest public database of curated lipid structures with >45,000 unique lipid entries classified under the LIPID MAPS classification system (fatty acyls, glycerolipids, glycerophospholipids, sphingolipids, sterols, prenols, saccharolipids, polyketides). Provides mass spectrometry reference data (predicted m/z), systematic nomenclature, and 2D/3D structure downloads. Essential resource for lipidomics, metabolomics, and membrane biology research. | Downloads, API limited | |
| βοΈ MetaCyc | Comprehensive curated database of >3,000 experimentally elucidated metabolic pathways from all domains of life, with >31,000 reactions and >20,000 compounds. MetaCyc pathways are manually reconstructed from primary literature with detailed enzyme information, reaction mechanisms, pathway variants, and organism-specific evidence. The gold-standard reference for metabolic pathway annotation, genome-scale metabolic model reconstruction, and comparative metabolomics. | Downloads, API via Pathway Tools π² | |
| βοΈ Reactome | Curated open-source pathway database covering >2,500 human pathways encompassing >13,000 reactions organized hierarchically from molecular events to biological processes. Provides manually curated reaction details with physical entities (proteins, chemicals, complexes), subcellular localization, literature evidence, and orthology-based pathway projections to 20+ other species. Primary resource for pathway enrichment analysis, reactome-based functional interpretation of omics data, and systems biology modeling. | π€ | ContentService API, Downloads |
| βοΈ Rhea | Curated database of >13,000 biochemical reactions with precisely defined reaction participants from ChEBI, stoichiometric coefficients, and directionality (bidirectional annotations). Each reaction is manually annotated with accurate chemical structures for substrates and products, enzyme catalysis (EC number), and cross-references to KEGG, MetaCyc, Reactome, and pathway databases. The gold-standard reference for enzyme-reaction mapping and building chemically accurate metabolic networks. | π€ | REST API, Downloads |
| βοΈ SABIO-RK | Curated database of biochemical reaction kinetics containing >100,000 kinetic parameter entries manually extracted from scientific literature, with standardized rate equations, reaction conditions (pH, temperature, buffer), and organism information. Provides SBML export for direct integration into computational metabolic models and supports parameter search by enzyme, pathway, or organism. Key resource for building kinetic models of metabolic networks and constraint-based modeling approaches. | π€ | REST API, Downloads |
| Name | Description | Agent | Links |
|---|---|---|---|
| π BindingDB | Public database of >2.6 million measured binding affinities (K_i, K_d, IC_50, EC_50) for protein-ligand and protein-protein interactions, spanning >8,500 protein targets and >1 million compounds. Provides experimentally measured binding data from multiple assay types with standardized units and cross-references to PDB structures, PubMed, and ChEBI/CHEMBL compound identifiers. Essential resource for training machine learning models for drug-target interaction prediction and virtual screening. | π€ | REST API, Downloads |
| π ChEBI | Chemical Entities of Biological Interest β EMBL-EBI dictionary of >190,000 small chemical compounds (metabolites, drugs, natural products) with manually curated annotation of biological roles, applications, and chemical classifications. Provides a structured ontology (ChEBI ontology) organizing compounds by role, structure, and source with >100,000 parent-child relationship links. The authoritative chemical ontology for linking compounds to biological function in bioinformatics databases. | π€ | REST API, FTP |
| π ChEMBL | Manually curated database of >2.3 million bioactive molecules with >20 million bioactivity measurements against >15,000 targets (proteins, nucleic acids, cell-based assays), extracted from >80,000 publications. Provides standardized activity data (IC_50, K_i, K_d, EC_50) with compound structures, target information, and ADMET data in a relational schema optimized for drug discovery. Primary benchmark resource for drug-target interaction prediction, virtual screening, and chemogenomic machine learning. | π€ | REST API, Downloads |
| π DGIdb | Drug-Gene Interaction database integrating >40,000 drug-gene interactions from 30+ sources including FDA labels, ClinicalTrials.gov, PharmGKB, DrugBank, and biomedical literature. Provides gene-level drug interaction summaries with clinically actionable variant annotations, druggable gene categories (kinases, ion channels, GPCRs), and search by gene, drug, or interaction type. Key resource for identifying drug repurposing opportunities and interpreting genomic variants in a clinical context. | π€ | GraphQL API, Downloads |
| π Drug Repurposing Hub | Broad Institute's curated collection of >6,000 approved and investigational drugs annotated with mechanism of action, therapeutic indications, clinical status, and chemical properties, optimized for systematic drug repurposing screening. Includes >1,200 FDA-approved drugs and >2,000 clinical-stage compounds with standardized compound handling protocols and screening data across >100 disease cell lines. Foundational resource for drug repurposing discovery and phenotypic screening campaigns. | π€ | Downloads, API via CLUE π |
| π DrugBank | Comprehensive drug knowledgebase containing >500,000 drug entries including FDA-approved small molecule drugs, biologics, nutraceuticals, and experimental compounds with detailed target, enzyme, transporter, and carrier information. Provides drug-drug interactions, metabolic pathways, ADMET data, pharmacoeconomic information, and drug-target-pathway-disease network associations. The most authoritative open resource for drug-centric knowledge, widely used for drug repurposing, mechanism prediction, and pharmaceutical AI. | Downloads (free for academic use), API restricted π² | |
| π Open Targets Platform | Comprehensive target-disease association evidence from >20 integrated data sources (GWAS, eQTL, ChIP-seq, CRISPR screens, animal models, literature), with >500,000 target-disease associations scored using an evidence-based framework. Provides Locus-to-Gene (L2G) scores linking GWAS loci to causal genes, drug mechanism-of-action annotations, and disease ontology mapping. Leading platform for AI-driven target identification, drug repurposing, and target prioritization in drug discovery. | π€ | GraphQL API, Downloads |
| π PDBbind | Curated database of experimentally measured binding affinities for protein-ligand complexes from the PDB, containing >23,000 entries (v2021) with K_i, K_d, and IC_50 values linked to 3D structural data. Provides refined sets (core, refined, general) benchmarked for structure-based drug design method evaluation. The standard benchmark for evaluating protein-ligand docking, scoring functions, and structure-based binding affinity prediction. | Downloads (registration required) | |
| π PubChem | World's largest public chemistry database with >115 million compounds, >300 million bioactivity test results, and >30 million substance descriptions aggregated from 850+ data sources. Provides standardized chemical structures (2D/3D), physicochemical properties, safety and toxicity data, literature references, and patent information through the flexible PUG REST API (supporting various input/output formats including SDF, JSON, XML, CSV). The primary public chemistry reference for drug discovery, chemical biology, and cheminformatics AI applications. | π€ | PUG REST API, FTP |
| π SMPDB | Small Molecule Pathway Database with >30,000 interactive metabolic, signaling, and disease pathway diagrams designed for human, with a subset for model organisms. Each pathway includes metabolite structures, protein annotations, drug targets, and disease associations linked to HMDB, DrugBank, and UniProt entries. Valuable resource for visualizing and interpreting metabolomics and pathway analysis results alongside human disease mechanisms. | Downloads, no API | |
| π STITCH | Search Tool for Interacting Chemicals integrating known and predicted chemical-protein interactions for >500,000 chemicals across >9 million proteins from 2,000+ organisms. Combines evidence from experimental binding data, pathway knowledge, text-mining of PubMed, and in silico interaction predictions with confidence scoring. Connected to the STRING network framework for context-based exploration of chemical effects on cellular networks. | Downloads, API via STRING | |
| π SuperDRUG2 | Comprehensive database of approved and marketed drugs worldwide, covering >11,000 active pharmaceutical ingredients with regulatory information (FDA, EMA, PMDA), chemical structures, therapeutic targets, metabolism, and side effect profiles. Provides drug-drug interaction data, pharmacokinetic parameters, and cross-references to DrugBank, PubChem, and KEGG. Useful resource for pharmaceutical research, drug safety assessment, and drug repurposing studies starting from approved molecules. | β | Web-based search and download |
| π ZINC | Free database of >1.5 billion commercially available compounds for virtual screening, with >230 million in ready-to-dock 3D formats (mol2, pdbqt) from >200 vendor catalogs. Provides multiple subsets (drug-like, lead-like, fragment-like) with physicochemical property filters, pre-calculated conformations, and varying protonation states. The standard resource for structure-based virtual screening, molecular docking, and compound library design in academic drug discovery. | π€ | REST API, Downloads |
| Name | Description | Agent | Links |
|---|---|---|---|
| π₯ cBioPortal | Open-source platform for interactive exploration of >390 cancer genomics studies encompassing >100,000 tumor samples, integrating somatic mutations, copy-number alterations, mRNA/protein expression, DNA methylation, and pathway-level changes. Provides signature visualizations including OncoPrint, mutual exclusivity analysis, survival curves, and cohort comparison across studies. Primary platform for exploring large-scale cancer genomics data (TCGA, AACR GENIE, MSK-IMPACT) and generating data-driven cancer hypotheses. | π€ | REST API, Downloads |
| π₯ ClinicalTrials.gov | NIH registry and results database of >450,000 clinical studies conducted in 220+ countries, with structured information on study design, interventions, eligibility criteria, outcomes, and results for both interventional and observational trials. Provides a REST API v2 for programmatic search and retrieval, supporting complex queries by condition, intervention, location, sponsor, and study phase. The global authoritative registry for clinical trial transparency, meta-analysis, and clinical research informatics. | π€ | REST API v2, Downloads |
| π₯ ClinVar | NCBI database of >3 million submitted human genetic variant interpretations with standardized clinical significance classifications (pathogenic, benign, uncertain significance) following ACMG/AMP guidelines. Uses a star-rating review status system (0-4 stars) based on the number of submitters and review criteria, enabling confidence-based filtering for clinical decision support. The gold-standard resource for clinical variant interpretation, hereditary disease diagnostics, and training genomic AI models. | π€ | E-Utilities API, FTP |
| π₯ COSMIC | Catalogue of Somatic Mutations in Cancer β the most comprehensive somatic mutation database with >25 million coding mutations curated from >40,000 publications and large-scale cancer genome screens. Provides the Cancer Gene Census (CGC) of >700 genes with causative roles in cancer (Tier 1 and Tier 2 classifications), mutational signatures, and drug resistance mutation annotations. Essential resource for cancer genomics driver discovery, mutational signature analysis, and precision oncology. | Downloads (registration + license required) | |
| π₯ DepMap | Cancer Dependency Map β genome-scale CRISPR knockout (Achilles) and RNAi screens across >1,000 cancer cell lines, integrating multi-omics characterization from CCLE (gene expression, mutations, copy-number, methylation, protein) and drug sensitivity data from PRISM. Provides CERES gene effect scores quantifying essentiality of each gene in each cell line, enabling identification of cancer-specific vulnerabilities. Primary resource for cancer vulnerability discovery, synthetic lethality prediction, and precision oncology target identification. | π€ | Downloads + Portal, API via data files |
| π₯ DisGeNET | Knowledge platform integrating >1.1 million gene-disease associations covering >30,000 diseases and >21,000 genes, with evidence scores aggregated from expert-curated databases (CTD, ClinGen, Orphanet) and text-mined literature. Provides variant-disease links (VDA), gene-disease association scores (GDA), and disease-disease similarity based on shared genetic associations, with versioned downloads for reproducible analysis. Key resource for rare disease gene discovery, disease gene prioritization, and investigating disease mechanisms. | π€ | REST API, Downloads π |
| π₯ HPO | Human Phenotype Ontology β standardized vocabulary of >20,000 phenotypic abnormality terms with hierarchical structure (subclass/parent-child relationships) for describing human disease manifestations. Provides >200,000 phenotype-to-gene annotations from OMIM, Orphanet, and DECIPHER, enabling phenotype-driven exome/genome prioritization. Foundational resource for computational phenotyping, rare disease diagnosis via phenotype matching, and genotype-phenotype association studies. | π€ | REST API, Downloads |
| π₯ ICGC | International Cancer Genome Consortium β an international collaboration aggregating genomic data from >25,000 cancer cases across 50+ tumor types and 17+ contributing projects worldwide. Provides standardized somatic mutation calls (SNVs, indels, structural variants, copy-number alterations), gene expression, and methylation data with controlled access for sensitive data. Complementary to TCGA, ICGC adds diverse population representation and additional rare tumor type coverage for global cancer genomics. | Downloads, API access limited | |
| π₯ MIMIC-IV | Medical Information Mart for Intensive Care β de-identified electronic health record data from >300,000 ICU patients and >500,000 hospital admissions at Beth Israel Deaconess Medical Center (2008-2019). Provides detailed clinical data including vital signs, laboratory measurements, medications, diagnoses, procedures, fluid balance, and clinical notes with timestamped events. The premier open-access critical care database for machine learning in health care, clinical prediction modeling, and health services research. | Downloads via PhysioNet (credentialed access + CITI training required) | |
| π₯ Monarch Initiative | Integrative platform connecting genotypes, phenotypes, and diseases across human and >800 model organisms (mouse, zebrafish, fly, worm, yeast) using standardized ontologies (HPO, MP, ZP, GO). Provides cross-species phenotype matching for rare disease diagnostics, gene-phenotype association discovery, and a knowledge graph linking genes, diseases, phenotypes, and anatomical entities. Key resource for rare disease gene discovery through cross-species translational genomics. | π€ | REST API, Downloads |
| π₯ OMIM | Online Mendelian Inheritance in Man β the authoritative, expert-curated catalog of >26,000 human gene-phenotype relationships covering >16,000 genes and all known Mendelian disorders (over 8,000 phenotypic entries). Each entry provides a detailed synopsis of clinical features, genetic heterogeneity, inheritance patterns, molecular genetics, and allelic variants with comprehensive literature reviews. The gold-standard reference for clinical genetics, Mendelian disease gene identification, and variant interpretation in rare disease diagnostics. | API (limited free tier), downloads restricted | |
| π₯ Open Targets Genetics | Portal integrating GWAS catalog data, UK Biobank, FinnGen, and other large-scale genetic studies with functional genomics (QTL, chromatin interaction, variant annotation) for identifying causal genes and potential drug targets. Provides variant-to-gene prioritization (L2G), colocalization analyses (eQTL-GWAS), and fine-mapping results across thousands of disease-associated loci. Essential resource for translating genetic associations into actionable drug target hypotheses using statistical genetics. | π€ | GraphQL API, Downloads |
| π₯ Orphanet | Portal for rare diseases and orphan drugs β the authoritative reference for >10,000 rare diseases with unique Orpha codes, classifications, epidemiological data (prevalence, incidence), clinical signs, and causal gene information. Provides an inventory of orphan drugs with marketing authorization status, expert centers directory, and emergency medical guidelines for rare disease management. Primary resource for rare disease coding, clinical research, health policy, and drug repurposing analysis. | Downloads, no public API | |
| π₯ TCGA/GDC | NCI Genomic Data Commons β unified repository and data sharing platform hosting genomic, transcriptomic, epigenomic, and clinical data from TCGA (11,000+ tumors across 33 cancer types) and other NCI cancer programs (TARGET, CPTAC). Provides standardized data processing pipelines (harmonization) for somatic mutation calling, gene expression quantification, and copy-number analysis with controlled access for clinical data. The foundational resource for cancer genomics, biomarker discovery, and AI-based cancer diagnosis and prognosis models. | π€ | REST API, Downloads |
| Name | Description | Agent | Links |
|---|---|---|---|
| π BioGRID | Curated biological interaction database containing >2.5 million protein-protein, genetic, and chemical interactions from 80+ model organisms (yeast, human, fly, worm, mouse, Arabidopsis, and others). Manually curated from both low-throughput and high-throughput studies with standardized interaction types (physical association, direct interaction, co-localization, genetic interaction). Primary resource for network biology, protein complex mapping, and training protein interaction prediction models. | π€ | REST API, Downloads |
| π CORUM | Comprehensive resource of manually annotated mammalian protein complexes, containing >5,000 protein complexes from human, mouse, and rat mainly, each curated from the primary literature with complex composition, function, cellular localization, and disease associations. Provides evidence for each complex member (stoichiometry, purification method, identification method) with cross-references to UniProt and PubMed. Essential reference for studying protein complex assembly, complex-based functional annotation, and systems biology modeling. | Downloads, no API | |
| π DRKG | Drug Repurposing Knowledge Graph β comprehensive biomedical knowledge graph integrating 177,000+ entities (drugs, diseases, genes, pathways, side effects) and >2.3 million relationships from 30+ public databases (DrugBank, CTD, STRING, IntAct, GNBR, and others). Specifically designed for knowledge graph embedding and link prediction for drug repurposing, providing pre-split train/test sets for reproducible evaluation. Popular benchmark for biomedical knowledge graph completion and computational drug repurposing. | Downloads, no public API | |
| π Hetionet | Heterogeneous knowledge graph (hetnet) integrating >47,000 nodes of 11 types (genes, compounds, diseases, anatomies, pathways, biological processes, molecular functions, cellular components, pharmacologic classes, side effects, symptoms) connected by >2.25 million relationships of 24 types. Built from 29 public databases and manually curated with consistent edge weights and directionality. Pioneering resource for biomedical knowledge graph analytics, drug repurposing via network propagation, and graph neural network benchmarks. | π€ | Downloads, API via Neo4j |
| π HIPPIE | Human Integrated Protein-Protein Interaction rEference β scored human protein-protein interaction network with >300,000 PPIs between >20,000 human proteins, integrating evidence from curated databases and high-throughput experiments. Each interaction is assigned a confidence score based on the number and reliability of supporting experiments, with tissue-specific and disease-specific context annotations. Useful for human-specific PPI network analysis, disease module identification, and functional genomics interpretation. | Downloads, web-based API | |
| π IID | Integrated Interactions Database providing tissue-specific and species-specific protein-protein interaction networks for human and 8 model organisms (mouse, rat, zebrafish, fly, worm, yeast, Arabidopsis, E. coli). Integrates data from 10+ primary interaction databases with tissue expression filtering to compute >20 million context-specific interactions across >100 tissues. Enables tissue-aware network biology and studying cell-type-specific interactome rewiring in disease. | Downloads, no public API | |
| π IntAct | EMBL-EBI molecular interaction database and IMEx consortium member, providing >1.2 million curated binary interactions covering protein-protein, protein-DNA, protein-RNA, and protein-small molecule interactions. Uses the standardized PSI-MI XML format for data exchange and the MI ontology for controlled vocabulary annotation of interaction detection methods, participant roles, and interaction types. Key resource for training AI models for protein interaction prediction and network-based functional annotation. | π€ | REST API, FTP |
| π PathwayCommons | Aggregated pathway and interaction data from 22 public pathway databases (Reactome, KEGG, WikiPathways, PID, NCI-Nature, PANTHER, and others) in BioPAX and Simple Interaction Format (SIF). Provides a graph search API for retrieving interactions and pathways across all integrated sources, supporting queries by protein, compound, or pathway. Key integration hub for pathway-level analysis and building unified interaction networks from heterogeneous sources. | π€ | REST API, Downloads |
| π PrimeKG | Precision Medicine Knowledge Graph β multimodal knowledge graph integrating 20+ biomedical resources covering diseases, drugs, proteins, genes, biological processes, pathways, and side effects into >130,000 nodes and 4 million relationships. Designed for precision medicine AI, supporting drug-disease prediction, drug repurposing, and patient stratification via graph machine learning. Provides pre-computed graph embeddings and links to original source databases for validation. | π€ | Downloads, API via Neo4j |
| π STRING | Search Tool for the Retrieval of Interacting Genes/Proteins β the largest protein-protein interaction network covering >67 million proteins from >14,000 organisms, with >20 billion predicted and known interactions. Integrates evidence from experimental data, curated databases, text mining of >100 million PubMed abstracts, co-expression analysis, and interolog (cross-species) transfer, all scored with a combined confidence framework. Provides functional enrichment analysis (GO, KEGG, Reactome, UniProt keywords) for protein lists and network clustering. | π€ | REST API, Downloads |
| π TRRUST | Transcriptional Regulatory Relationships Unraveled by Sentence-based Text mining β manually curated database of >8,000 transcription factor (TF)-target gene interactions in human and mouse, extracted from biomedical literature using sentence-based text mining with expert validation. Covers 800+ human TFs and 400+ mouse TFs with information on regulatory direction (activation/repression), literature support, and TF-TF regulatory networks. Key resource for reconstructing gene regulatory networks and studying transcriptional regulation. | Downloads, no API | |
| π WikiPathways | Community-curated biological pathway database in a wiki framework, containing >3,200 pathways across 30+ species contributed and maintained by the scientific community. Provides pathways in multiple formats (GPML, BioPAX, SVG) with interactive pathway diagrams, gene product and metabolite cross-references, and pathway ontology annotations. Unique resource for capturing newly discovered pathways and model organism-specific pathways not yet covered by major pathway databases. | π€ | REST API, Downloads |
| Name | Description | Agent | Links |
|---|---|---|---|
| π ArrayExpress | EMBL-EBI functional genomics data archive (now merged into BioStudies) containing >70,000 experiments from microarray and high-throughput sequencing platforms across all domains of life. Provides standardized data in MAGE-TAB and ISA-Tab formats with MIAME/MINSEQE-compliant metadata, enabling reproducible reanalysis of functional genomics experiments. Main repository for ArrayExpress-deposited functional genomics data alongside GEO. | π€ | REST API, FTP |
| π Expression Atlas | EMBL-EBI gene expression resource covering >3,500 experiments across 50+ species, providing both baseline expression (gene expression levels across tissues/cell types) and differential expression (disease vs. normal, treatment vs. control). Integrates transcriptomics (RNA-seq, microarray) and proteomics data from ArrayExpress, all reprocessed through standardized analysis pipelines. Enables cross-species and cross-study comparison of gene expression with direct access to normalized expression values. | π€ | REST API, Downloads |
| π FANTOM5 | Functional Annotation of Mammalian Genome 5 β CAGE (Cap Analysis of Gene Expression)-based promoter-level expression atlas profiling >1,800 human and mouse primary cell types, tissues, and time-course samples. Provides genome-wide promoter usage, transcription start site (TSS) annotations, enhancer identification, and promoter-level expression quantification with the FANTOM5 promoterome. Unique resource for studying promoter usage dynamics, enhancer RNA biology, and cell-type-specific transcriptional regulation. | Downloads, no public API | |
| π GEO | NCBI Gene Expression Omnibus β the largest public functional genomics data repository with >5 million samples from >200,000 studies spanning microarray, RNA-seq, ChIP-seq, ATAC-seq, and other high-throughput platforms across all domains of life. Provides MIAME/MINSEQE-compliant metadata, processed expression matrices, and raw data files with programmatic access via the E-Utilities API. The primary repository for reanalyzing published expression data and training gene expression-related machine learning models. | π€ | E-Utilities API, FTP |
| π GTEx | Genotype-Tissue Expression project β tissue-specific gene expression and genetic regulation data from 948 donors (v8) across 54 non-diseased human tissue sites. Provides comprehensive eQTL (cis- and trans-), splice QTL (sQTL), and allele-specific expression (ASE) associations linking genetic variants to molecular phenotypes across tissues. Protected health data (individual-level genotypes, RNA-seq BAMs) accessed through dbGaP controlled tiers. Essential resource for understanding human gene regulation, tissue-specific disease mechanisms, and functional interpretation of GWAS loci. | π€ | REST API, Downloads |
| π HOCOMOCO | Comprehensive collection of >1,000 hand-curated human and mouse transcription factor binding models (position frequency matrices, PFMs) constructed from high-quality ChIP-seq datasets. Provides mononucleotide and dinucleotide PWM variants, motif enrichment analysis tools (HOCOMOCO Scanner), and TF binding site annotation across the human and mouse genomes. Key reference for TF binding site prediction, motif enrichment analysis in ChIP-seq/ATAC-seq peaks, and studying TF binding specificity. | π€ | Downloads, API via MEME Suite format |
| π JASPAR | Open-access database of >2,000 curated, non-redundant transcription factor binding profiles (PFMs, PWMs) across 6 taxonomic groups (vertebrates, insects, plants, nematodes, fungi, urochordates). Each profile is manually curated from published ChIP-seq, SELEX, and protein-binding microarray experiments with cell-type and method provenance information. The community-standard reference for TF binding site prediction, motif scanning with FIMO/MAST, and evaluating TF binding specificity in regulatory genomics. | π€ | REST API, Downloads |
| π mirBase | microRNA database (miRBase v22) β the primary online repository for published microRNA sequences and annotation, containing >38,000 hairpin precursors and >48,000 mature miRNA sequences from 271 organisms. Provides standardized nomenclature (miR- naming system), genomic coordinates, evidence for miRNA annotation (cloning, sequencing, expression), and predicted targets via TargetScan. The definitive reference for miRNA sequence information and miRNA family classification. | FTP downloads, no API | |
| π miRTarBase | Database of experimentally validated microRNA-target interactions with >500,000 entries validated by reporter assays, Western blot, qRT-PCR, microarrays, and NGS-based crosslinking immunoprecipitation (HITS-CLIP, PAR-CLIP). Covers >11,000 miRNAs and >400,000 target genes across 28 species, with functional annotations including miRNA disease associations and functional GTEx-type tissue expression. Essential resource for miRNA biology, miRNA-based therapeutic target identification, and training miRNA-target prediction models. | Downloads, no API | |
| π RegNetwork | Repository of integrated transcriptional (TF-target) and post-transcriptional (miRNA-target, TF-miRNA) regulatory relationships in human and mouse, aggregated from multiple curated databases (TRRUST, miRTarBase, TransmiR, and others). Provides downloadable regulatory networks in standard formats for network analysis and visualization. Useful resource for constructing gene regulatory networks and studying multi-layer regulatory interactions in human and mouse. | β | Web-based query and download |
| π Roadmap Epigenomics | NIH Roadmap Epigenomics Project providing comprehensive reference epigenomic maps (histone modifications, DNA methylation, chromatin accessibility, RNA-seq) across >100 human tissues and primary cell types, including 25 core epigenomes (embryonic stem cells, blood, brain, heart, liver, and others). Generated using standardized experimental and computational pipelines for ChIP-seq, WGBS, DNase-seq, and RNA-seq. Foundational resource for studying cell-type-specific regulatory elements, enhancer-gene links, and epigenetic mechanisms in human biology and disease. | Downloads via GEO/SRA, no unified API | |
| π TRANSFAC | Curated database of eukaryotic transcription factors, their DNA binding sites (position weight matrices and genomic binding sites), and regulated target genes, maintained and regularly updated by professional curators from the literature. Covers human, mouse, rat, and other model organisms with tissue-specific expression patterns for TFs and binding site annotations. Widely used for promoter analysis, TF binding site prediction, and gene regulatory network reconstruction in the pharmaceutical and biotechnology industry. | π² | Commercial license required for full access; limited public version at PATCH |
See CONTRIBUTING.md for guidelines on adding new databases.
Detailed API endpoints and curl examples for Agent-Ready databases are in api-reference.md.
The database list and descriptions in this repository are dedicated to the public domain under CC0 1.0 Universal. All factual information about databases (names, URLs, access methods) is not subject to copyright.