Open development of genomic language models — data, modeling, and evaluation.
Inspired by Marin.
- 2026-08-03 — Blog post — A 1B standard Transformer rivals Evo 2 40B on variant effect prediction.
- 2026-05-26 — Poster — Data curation strategies for genomic language models.
These documents synthesize MarinDNA's current answers and help organize future experiments.
- Bidirectionality
- Context size
- Data mixing
- Evolutionary timescales
- Latent biological features
- Sequence-to-function modeling
- Species conditioning
- Tokenization
- Training regions
- Models and datasets on Hugging Face
- Variant effect prediction leaderboard
- Interactive sequence explorer
- Model inference and BRCA1 variant effect prediction notebook
Join the Marin Discord; MarinDNA discussion happens in the #dna channel.
If you find datasets, models, or experiments from this repo useful, please cite:
MarinDNA: open development of genomic language models. Open Athena, 2026. https://github.com/Open-Athena/marin-dna
BibTeX:
@misc{marin-dna,
title = {MarinDNA: open development of genomic language models},
author = {{Open Athena}},
year = {2026},
url = {https://github.com/Open-Athena/marin-dna},
}