Variant in. Ranked CRISPR designs out.
Design and compare SpCas9, base-editing, and prime-editing strategies in one uncertainty-aware, population-aware workflow.
Warning
AlleleForge is a research tool, not a medical device. Its designs and off-target nominations are computational hypotheses that require experimental validation.
CRISPR design often means moving a variant through separate tools for guide selection, efficiency scoring, outcome prediction, and off-target analysis. AlleleForge puts that work behind one interface.
Give it a variant and a reference genome. It determines which editing chemistries can make the change, generates candidates, predicts their outcomes, evaluates off-targets, and returns one ranked menu with the evidence needed to compare them.
- Compare editing strategies together. Evaluate SpCas9 nuclease, ABE/CBE base editors, and prime editors through the same pipeline.
- See likely outcomes alongside guide scores. Candidates include predicted intended edits, indels, bystanders, and other byproducts relevant to the chemistry.
- Account for human variation. Add population frequencies, phased haplotypes, or a patient VCF to find off-target sites absent from the reference genome. Runs without those inputs are clearly labeled reference-only.
- Keep uncertainty visible. Predictions carry intervals, distribution checks, and an explicit calibration status.
- Reproduce and audit results. Reports record the reference, models, datasets, settings, seed, licenses, and content hashes used to produce them.
- Move from design to review. Export JSON, TSV, Parquet, HTML, PDF, and cloning-ready oligos from the same core library.
AlleleForge is available as a Python library, the aforge CLI, a web API and browser
interface, and the CRISPR-Bench evaluation harness.
AlleleForge requires Python 3.11 or 3.12. The project is currently an active alpha, so install it from source:
git clone https://github.com/clay-good/alleleforge.git
cd alleleforge
python3 -m venv .venv
source .venv/bin/activate
make installTurn a genomic variant into a ranked HTML report:
aforge design 'chr11:5227002:A>T' \
--reference-fasta hg38.fa \
--intent correct \
--format html \
--out report.htmlAdd population-aware off-target analysis by supplying allele frequencies:
aforge design 'chr11:5227002:A>T' \
--reference-fasta hg38.fa \
--gnomad gnomad.sites.tsv.gz \
--populations afr,eur,eas \
--format html \
--out report.htmlFetch a pinned dataset snapshot. The command prints its verified cache path; ClinVar can
be passed to --clinvar, while GENCODE's GTF loads directly through GeneModels:
CLINVAR_PATH=$(aforge data fetch clinvar)
aforge data fetch gencode
aforge resolve VCV000012345 --reference-fasta hg38.fa --clinvar "$CLINVAR_PATH" --jsonaforge data status re-hashes cached and bundled artifacts before calling them available;
a corrupt file is reported as cached but unverified, with a refresh remedy.
The path printed by aforge data fetch clinvar can be passed directly to
aforge resolve, aforge design, or aforge batch with --clinvar.
The same pipeline is available in Python:
from alleleforge.design import design
from alleleforge.genome import ReferenceGenome
with ReferenceGenome("hg38.fa", build="hg38") as reference:
menu = design("chr11:5227002:A>T", reference=reference)
print(menu.model_dump_json(indent=2))Run aforge --help for variant resolution, cohort design, standalone off-target
search, result verification, model and dataset inspection, caching, and benchmarking.
| Input | Output |
|---|---|
| Genomic coordinates, VCF records, genomic HGVS, ClinVar accessions, dbSNP IDs, or raw target sequences | Ranked candidates across every applicable editing chemistry |
| A local reference FASTA | Guides, pegRNAs, donor designs, and cloning oligos |
| Optional gnomAD-style frequencies, haplotypes, and patient variants | Reference, population, haplotype, and patient-specific off-target findings |
| Optional trained models and chromatin tracks | Efficiency and outcome predictions with provenance and uncertainty |
Reports summarize the top 3 predicted outcome alleles per candidate by default. Use
--top-alleles N (or top_alleles in the web request) to change that summary; the
ranked-menu output always retains the complete outcome spectrum.
ClinVar and GENCODE have checksum-pinned snapshots available through aforge data fetch;
dbSNP, coding/protein HGVS, trained models, and the remaining external annotations require
their documented optional dependencies or data sources. The default pipeline uses
transparent, weight-free baselines. A data fetch/data refresh command is explicit
consent for that artifact download; automatic downloads and VEP lookups remain separately
consent-gated.
For commercial work, set ALLELEFORGE_MODEL_USE=commercial. The license gate refuses
trained models that do not permit the use you declare.
The optional Rust extension accelerates the off-target hot path while keeping a
byte-identical Python fallback. make native builds the release wheel and runs the
entire suite against it. Performance is reported, not gated: hardware and workloads
matter, and an exact native algorithm can still be slower end to end.
Recorded on September 28, 2026, with Python 3.11.16 on arm64 macOS using the release build and the synthetic workloads in the benchmark harness:
| Kernel | Current hot path | Recorded result |
|---|---|---|
bwt |
Opt-in persistent FM-index query | 0.03x at 300 kb and 1 Mb, so slower than the default linear scan |
kmer |
Opt-in exact seed prefilter | 6.1x for seed lookup, but 0.2x for the selective end-to-end scan |
haplotype |
Population haplotype materialization | 4.0x |
The same run measured the whole native strand scan at 1.4x, bulged alignment at
10.8x, and resolved-base counting at 4.0x. Re-run
python scripts/native_speedup.py on the deployment hardware before using these
numbers for capacity planning. The FM-index and k-mer prefilter stay opt-in because
their end-to-end measurements do not beat the default path.
- Documentation
- Population-aware off-target analysis
- Uncertainty contract
- Python API
- CLI reference
- Data and provenance
- Runnable notebooks
- Current v1.0 work
make ci # includes artifact, SBOM, and immutable workflow-action checks
make image # build and import-smoke the deployment image (requires Docker)CI actions and the deployment base image are content-pinned, and the image resolves
its Python runtime from constraints/container.txt;
Dependabot proposes reviewable updates instead of allowing upstream releases to change
a green commit. Published images carry max-level build provenance and an OCI-attached
SBOM; GitHub Releases carry SHA256SUMS for the wheel, sdist, and wheel SBOM. The served
container runs as unprivileged UID/GID 10001:10001 and advertises its
/api/health liveness check to Docker and Compose. The default Compose service also
drops all Linux capabilities and makes its root filesystem read-only. The runtime
contains the installed package only, not a second source checkout, and a deny-by-default
build context excludes local reference data.
CI's visible advisory job audits both the development environment and the exact container dependency snapshot; a finding stays visible without blocking unrelated work until the v1.0 policy makes that audit mandatory.
Contributions are welcome. See CONTRIBUTING.md, and report security issues through the private process in SECURITY.md.
AlleleForge produces research hypotheses, not clinical decisions. Computational off-target analysis does not replace experimental validation such as GUIDE-seq, CHANGE-seq, or targeted amplicon sequencing. Review model cards and dataset provenance before acting on a result, especially when a prediction is uncalibrated or outside its training distribution.
AlleleForge runs locally by default and has no telemetry. Optional external services state what data they send and require separate consent.
AlleleForge is released under the MIT License. Wrapped models and tools keep their upstream licenses, which the model and data registries record and enforce.
If you use AlleleForge in research, cite it using CITATION.cff.