ESMFold2 on AtomWorks structures, with native ESMFold2 feature semantics preserved exactly.
Not a rewrite. The published model is held intact, fed from and returned to AtomWorks structures — then checked to be the same model: the same features, and the same outputs under deterministic kernels. An optional Foundry integration adds training, configuration and checkpointing on Foundry's trainer, without changing the native model.
AtomWorks AtomArray ──adapter──> ESMFold2 (unmodified) ──adapter──> AtomWorks AtomArray
│
└── feature parity, exact, on CPU, no weights
| part | state |
|---|---|
| AtomWorks ↔ ESMFold2 — adapter and its reverse, AtomWorks pipeline, supervision labels | done; feature parity exact on 5 fixtures |
| Inference — model wrapper, engine, CLI, a provenance record per fold | done; output parity exact on a GPU under deterministic kernels |
| Foundry integration (optional) — trainer, Hydra configs | contracts verified against the pinned Foundry checkout, the runtime ones also against rc-foundry 0.2.0; the objective is the caller's (docs/05) |
The core milestone is met: for monomer, multimer, metal, cofactor and modified-residue systems, the input the adapter derives from a structure featurizes to the same 29 tensors as an independently frozen reference input. Featurization is a pure function at a fixed seed, so that equality is exact. (Only a SMILES ligand makes the seed matter: its conformer is embedded at call time.)
The model's structure head is a diffusion sampler, and on a GPU with torch's deterministic algorithms off a seeded fold is not repeatable, so identical features alone say that the two paths condition the model identically. Under deterministic kernels they say more, and it is confirmed on a GPU: the two paths return identical structures, every coordinate and confidence value equal (docs/02).
pip install -e . # esm and atomworks resolve like any dependency
pip install -e ".[foundry]" # optional: training through Foundry (Python 3.12)
esmfold2-atomworks doctor # imports, weights; parity given AtomWorks' test data
esmfold2-atomworks parity structures/*.cif # survey a corpus
esmfold2-atomworks fold input.cif --out-dir runs/ # intended for a GPUfrom atomworks.io import parse
from esmfold2_atomworks.model.esmfold2 import AtomWorksESMFold2
parsed = parse("2hhb.cif.gz")
model = AtomWorksESMFold2()
structure, result = model.fold_atom_array(
parsed["asym_unit"][0], chain_info=parsed["chain_info"]
)The adapter fails loudly on the cases that otherwise produce a confident prediction of the wrong molecule:
- a ligand labelled
LIG/UNL/UNK/UNXis refused, because those are real CCD codes and what a model writes on anything it was handed as SMILES; - a declared
expected_formulathat disagrees with the heavy atoms present is an error; - a sequence is taken from the full entity record, not from the residues that happen to be modelled, so unresolved loops are not silently deleted;
- a non-standard residue is declared by CCD code, not folded as its parent;
- a chain kind, sequence override, MSA or ligand declaration that does not bind to exactly the chain it names — a typo, the wrong kind of chain, an alignment built for another sequence — is an error, not a no-op;
- a chain that holds more than one molecule — a protein, its ligands and its
waters under one author chain id, or the copies of a built assembly under one
chain_id— is an error, whether its annotations say so or a declared kind is contradicted by the CCD; - a fold with no ESMC backbone attached and no LM hidden states supplied is refused: the native model would fold anyway, without the LM prior the checkpoint was trained with;
- a chain that cannot be expressed, a covalent bond that cannot be placed, a chain kind that would have to be guessed, or a modification whose position is unknown each raises by default — accepted only by name, on every path, and recorded per output when accepted.
Each of these has a decision ID in docs/00_SCOPE.md and a test; the strictness policy is laid out in docs/01.
src/esmfold2_atomworks/
data/ the adapter, the reverse adapter, the AtomWorks pipeline, labels
model/ AtomWorksESMFold2 — holds the native module
inference/ engine: one resident model, many structures
parity/ feature and output comparison
metrics.py paths.py doctor.py cli.py
training/ optional Foundry integration: FabricTrainer subclass
configs/ optional Foundry integration: Hydra configs
tests/ parity and adapter tests; CPU-only by default
docs/ the design record — read docs/README.md first
Nothing outside training/ and configs/ imports Foundry, and
tests/test_core_boundary.py holds it to that. The optional integration is
described in docs/06.
pip install -e . is all the package needs: esm and atomworks are ordinary
dependencies, resolved from PyPI (Python ≥ 3.12; the foundry extra needs 3.12,
as rc-foundry does). The structure-based tests, feature parity among them,
read AtomWorks' test structures, which come with an atomworks checkout rather
than with the package: given those, the suite passes against the PyPI releases
of the upstreams, and without them those tests skip. See
docs/04.
The published parity results were produced in one pinned reference
environment: Python 3.14, torch 2.14, and the upstreams as source trees at the
revisions in reproducibility/UPSTREAM.lock. reproducibility/
defines it (uv sync --project reproducibility, then
source reproducibility/env.sh). It is not needed for ordinary use.
MIT — see LICENSE. This project wraps, but does not vendor, ESM (MIT) and AtomWorks / Foundry (BSD-3-Clause), each of which carries its own license and must be obtained separately.