Bioinformatics & Computational Biology | Python | NGS | Metagenomics | Proteomics | Machine Learning
I build reproducible computational workflows for biological data analysis, with interests spanning genomics, metagenomics, proteomics, machine learning, workflow automation, and research software practices.
- Programming & data analysis: Python, pandas, scikit-learn, matplotlib
- Bioinformatics: NGS, FASTQ processing, gene-expression analysis, metagenomics, proteomics
- Machine learning: classification, feature importance, model evaluation, explainability
- Workflow & reproducibility: Snakemake, SLURM/HPC, Git, GitHub Actions, pytest
- Data formats & tools: FASTQ, CSV/TSV, JSON, parquet, DIA-NN-style outputs
Reproducible metagenomics and machine-learning workflow with leakage-aware preprocessing, real output visualization, testing, GitHub Actions CI, Snakemake, and SLURM/HPC examples.
Gene-expression ML project using leakage-safe preprocessing, multiple classifiers, exported evaluation metrics, SHAP-based explainability, testing, CI, Snakemake, and SLURM.
Reproducible DIA-NN post-processing and QC workflow with CLI tools, validation, generated QC visualization, testing, GitHub Actions, Snakemake, and SLURM examples.
Tested Python CLI for paired-end FASTQ interleaving with gzip support, unequal-pair detection, optional mate-ID validation, synthetic examples, and CI.
I am especially interested in building analysis workflows that are clear, reproducible, testable, and suitable for collaborative scientific environments. My portfolio projects emphasize not only biological analysis, but also software quality through validation, automated tests, CI, workflow management, and HPC-aware execution.
- Bioinformatics and computational biology
- Genomics and transcriptomics
- Metagenomics and microbiome analysis
- Proteomics
- Machine learning for biological data
- Reproducible scientific computing
- Research workflow automation
My repositories are organized as public portfolio demonstrations. When original research data are unavailable, restricted, unpublished, or lab-owned, I use small example or synthetic datasets and clearly separate public reproducible code from broader project context.