Machine learning over measured protein fitness landscapes. Four related projects merged into one repository, each a self-contained package with its own tests, config, and commit history (imported via subtree merge).
cd active_learning_loop && pip install -e .
snakemake results/curves.png -j1 # fetches GB1 (FLIP), runs one seeded loopprotein-ml is the protein fitness ML repo: supervised and generative models over measured fitness landscapes (GB1, AAV), active learning, ESM embeddings, and diffusion. Sibling repos:
trust-tools (agent security
and evals), bio-qc (lab-data QC
pipelines), lab-informatics
(lab data plumbing and integrity),
llm-posttraining
(training-stage behavior work),
protein-ml (protein fitness
ML), and mol-ml (small-molecule
ML).
| Directory | What it does |
|---|---|
protein_diffusion/ |
Conditional DDPM over the measured GB1 fitness landscape. Reports memorization fraction and unmeasured-proposal handling as headline metrics. |
protein_stability_uncertainty/ |
Sequence-to-melting-point regression with split-conformal intervals, Mondrian binning, and sparse-bin fallback. |
protein_design_ops/ |
Closed design loop around upstream tools: ProteinMPNN generation, ESM-2 rescoring, ESMFold or Chai-1 structure prediction, and backbone self-consistency RMSD into a consensus report. |
active_learning_loop/ |
GP-UCB active learning over real fitness landscapes, with seeded random-baseline replication. |
Each package is independent. From its directory:
cd protein_diffusion && PYTHONPATH=src python -m pytest tests/ -qSame pattern for the other three. Each subdirectory retains its own
AGENTS.md with project-specific rules (data policy, claim standards,
verification commands), which still apply.
These four share datasets (FLIP / eLife GB1 mirrors), encoding utilities,
provenance manifests, and evaluation conventions. Merging them makes the
shared machinery visible in one place. A root-level parity test fails if
the vendored esm_cache.py copies drift apart.
