Skip to content

Latest commit

 

History

950 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AlleleForge

Variant in. Ranked CRISPR designs out.

Design and compare SpCas9, base-editing, and prime-editing strategies in one uncertainty-aware, population-aware workflow.

CI Python License: MIT

Warning

AlleleForge is a research tool, not a medical device. Its designs and off-target nominations are computational hypotheses that require experimental validation.

Why AlleleForge

CRISPR design often means moving a variant through separate tools for guide selection, efficiency scoring, outcome prediction, and off-target analysis. AlleleForge puts that work behind one interface.

Give it a variant and a reference genome. It determines which editing chemistries can make the change, generates candidates, predicts their outcomes, evaluates off-targets, and returns one ranked menu with the evidence needed to compare them.

  • Compare editing strategies together. Evaluate SpCas9 nuclease, ABE/CBE base editors, and prime editors through the same pipeline.
  • See likely outcomes alongside guide scores. Candidates include predicted intended edits, indels, bystanders, and other byproducts relevant to the chemistry.
  • Account for human variation. Add population frequencies, phased haplotypes, or a patient VCF to find off-target sites absent from the reference genome. Runs without those inputs are clearly labeled reference-only.
  • Keep uncertainty visible. Predictions carry intervals, distribution checks, and an explicit calibration status.
  • Reproduce and audit results. Reports record the reference, models, datasets, settings, seed, licenses, and content hashes used to produce them.
  • Move from design to review. Export JSON, TSV, Parquet, HTML, PDF, and cloning-ready oligos from the same core library.

AlleleForge is available as a Python library, the aforge CLI, a web API and browser interface, and the CRISPR-Bench evaluation harness.

Quick start

AlleleForge requires Python 3.11 or 3.12. The project is currently an active alpha, so install it from source:

git clone https://github.com/clay-good/alleleforge.git
cd alleleforge
python3 -m venv .venv
source .venv/bin/activate
make install

Turn a genomic variant into a ranked HTML report:

aforge design 'chr11:5227002:A>T' \
  --reference-fasta hg38.fa \
  --intent correct \
  --format html \
  --out report.html

Add population-aware off-target analysis by supplying allele frequencies:

aforge design 'chr11:5227002:A>T' \
  --reference-fasta hg38.fa \
  --gnomad gnomad.sites.tsv.gz \
  --populations afr,eur,eas \
  --format html \
  --out report.html

Fetch a pinned dataset snapshot. The command prints its verified cache path; ClinVar can be passed to --clinvar, while GENCODE's GTF loads directly through GeneModels:

CLINVAR_PATH=$(aforge data fetch clinvar)
aforge data fetch gencode
aforge resolve VCV000012345 --reference-fasta hg38.fa --clinvar "$CLINVAR_PATH" --json

aforge data status re-hashes cached and bundled artifacts before calling them available; a corrupt file is reported as cached but unverified, with a refresh remedy. The path printed by aforge data fetch clinvar can be passed directly to aforge resolve, aforge design, or aforge batch with --clinvar.

The same pipeline is available in Python:

from alleleforge.design import design
from alleleforge.genome import ReferenceGenome

with ReferenceGenome("hg38.fa", build="hg38") as reference:
    menu = design("chr11:5227002:A>T", reference=reference)

print(menu.model_dump_json(indent=2))

Run aforge --help for variant resolution, cohort design, standalone off-target search, result verification, model and dataset inspection, caching, and benchmarking.

What goes in and what comes out

Input Output
Genomic coordinates, VCF records, genomic HGVS, ClinVar accessions, dbSNP IDs, or raw target sequences Ranked candidates across every applicable editing chemistry
A local reference FASTA Guides, pegRNAs, donor designs, and cloning oligos
Optional gnomAD-style frequencies, haplotypes, and patient variants Reference, population, haplotype, and patient-specific off-target findings
Optional trained models and chromatin tracks Efficiency and outcome predictions with provenance and uncertainty

Reports summarize the top 3 predicted outcome alleles per candidate by default. Use --top-alleles N (or top_alleles in the web request) to change that summary; the ranked-menu output always retains the complete outcome spectrum.

ClinVar and GENCODE have checksum-pinned snapshots available through aforge data fetch; dbSNP, coding/protein HGVS, trained models, and the remaining external annotations require their documented optional dependencies or data sources. The default pipeline uses transparent, weight-free baselines. A data fetch/data refresh command is explicit consent for that artifact download; automatic downloads and VEP lookups remain separately consent-gated.

For commercial work, set ALLELEFORGE_MODEL_USE=commercial. The license gate refuses trained models that do not permit the use you declare.

Native acceleration

The optional Rust extension accelerates the off-target hot path while keeping a byte-identical Python fallback. make native builds the release wheel and runs the entire suite against it. Performance is reported, not gated: hardware and workloads matter, and an exact native algorithm can still be slower end to end.

Recorded on September 28, 2026, with Python 3.11.16 on arm64 macOS using the release build and the synthetic workloads in the benchmark harness:

Kernel Current hot path Recorded result
bwt Opt-in persistent FM-index query 0.03x at 300 kb and 1 Mb, so slower than the default linear scan
kmer Opt-in exact seed prefilter 6.1x for seed lookup, but 0.2x for the selective end-to-end scan
haplotype Population haplotype materialization 4.0x

The same run measured the whole native strand scan at 1.4x, bulged alignment at 10.8x, and resolved-base counting at 4.0x. Re-run python scripts/native_speedup.py on the deployment hardware before using these numbers for capacity planning. The FM-index and k-mer prefilter stay opt-in because their end-to-end measurements do not beat the default path.

Learn more

Development

make ci  # includes artifact, SBOM, and immutable workflow-action checks
make image  # build and import-smoke the deployment image (requires Docker)

CI actions and the deployment base image are content-pinned, and the image resolves its Python runtime from constraints/container.txt; Dependabot proposes reviewable updates instead of allowing upstream releases to change a green commit. Published images carry max-level build provenance and an OCI-attached SBOM; GitHub Releases carry SHA256SUMS for the wheel, sdist, and wheel SBOM. The served container runs as unprivileged UID/GID 10001:10001 and advertises its /api/health liveness check to Docker and Compose. The default Compose service also drops all Linux capabilities and makes its root filesystem read-only. The runtime contains the installed package only, not a second source checkout, and a deny-by-default build context excludes local reference data.

CI's visible advisory job audits both the development environment and the exact container dependency snapshot; a finding stays visible without blocking unrelated work until the v1.0 policy makes that audit mandatory.

Contributions are welcome. See CONTRIBUTING.md, and report security issues through the private process in SECURITY.md.

Responsible use

AlleleForge produces research hypotheses, not clinical decisions. Computational off-target analysis does not replace experimental validation such as GUIDE-seq, CHANGE-seq, or targeted amplicon sequencing. Review model cards and dataset provenance before acting on a result, especially when a prediction is uncalibrated or outside its training distribution.

AlleleForge runs locally by default and has no telemetry. Optional external services state what data they send and require separate consent.

License and citation

AlleleForge is released under the MIT License. Wrapped models and tools keep their upstream licenses, which the model and data registries record and enforce.

If you use AlleleForge in research, cite it using CITATION.cff.

About

A variant-first, multi-modality CRISPR design framework that unifies SpCas9 nuclease, base-editor, and prime-editor chemistries under a single typed interface to deliver ranked candidate edits complete with calibrated uncertainty intervals, predicted outcome distributions, and ancestry-stratified, population-aware off-target safety reporting.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages