A Curated List of Computational Biology Datasets Suitable for Machine Learning
-
Updated
Sep 2, 2026
A Curated List of Computational Biology Datasets Suitable for Machine Learning
🗺️ A repository of CNApy projects
Welcome to the Open Source Computational Biology Master's
Trace biological results from data and code to figures and scientific claims.
Bioinformatics analysis of adenovirus E1A proteins using Python, Biopython, LocalCIDER and AIUPred.
R integration package for the BioTrace GitHub Action. BioTrace traces biological results from data and code to figures and scientific claims.
Transformer-based system that designs novel protein structures for drug discovery and synthetic biology using evolutionary algorithms and deep learning.
Computational simulation of PLGA nanoparticle transport, drug release, and tumor response using finite difference methods in Python.
Curated list of AI tools for molecular modeling
Reproducible genomic prediction study comparing L1/L2-regularized artificial neural networks with GBLUP-ADE under a frozen cross-validation design, using canonical simulated-data workflows, workflowr, and renv.
Tool for accessing pre-processed single-cell spatial data.
ProSutra is an automated pipeline for protein structure prediction, validation, and report generation with CLI, Streamlit, and Electron interfaces.
Scrapes HDOCK job result pages, extracts the “Summary of the Top 10 Models” tables for any number of complexes, and builds a single-sheet Excel workbook with embedded download links to each job’s full results archive.
To associate your repository with the computational-biology-datasets topic, visit your repo's landing page and select "manage topics."