Research handoff: 2026-08-13 complete software/data handoff snapshot — navigation and governance; science authority remains the frozen 2026-08-01 bundle, and release/publication gates are still fail-closed.
Local mutation-state candidate reconstruction from single-molecule somatic-mutation co-occurrence in ONT long reads.
繁體中文版 → · Docs site → · Wiki → · How to run →
A tumour can contain several cellular populations carrying different mutation sets. Learning their order and cellular co-membership matters for resistance and progression, but those biological quantities are not directly observed by an unbarcoded bulk long read.
With short reads you only observe each locus' marginal variant allele frequency. Recovering the joint structure from those per-locus marginals alone, without linkage or additional model assumptions, is non-identifiable. ONT long reads change the problem: a single molecule can span several called somatic mutations at once. Their co-occurrence on the same physical molecule becomes a direct observation; cellular co-membership and lineage remain model-dependent inferences because the originating cell is not identified.
This repository implements that idea end to end — and, just as importantly, it is explicit about where the evidence runs out.
Five layers, bottom to top. Only some of them are currently runnable — see Status below, which is deliberately honest about what does and does not work.
| Layer | Component | What it does |
|---|---|---|
| Raw data | tumour / normal BAM, reference FASTA | ONT long reads carrying methylation tags |
| Upstream | Dorado · ClairS · LongPhase-S · SAVANA | basecalling + methylation, somatic SNV calling, haplotagging, copy number |
| Read-label representations | BAM HP/PS aux tags · HP/PS sidecar TSV | inter_sub_mod reads HP/PS from BAM aux tags; exact-PS/LongLineage use the sidecar. These are parallel, commit-scoped contracts, not a direct engine-to-engine runtime interface. |
| Engines | inter_sub_mod (C++17) · longlineage (C++17) |
per-region methylation/statistics · version-scoped per-read artefacts |
| Research/presentation | commit-pinned Python research solvers · validators/builders · standalone HTML | a Python solver may produce science only with its own inputs/receipt; publication builders and HTML present validated data without silently recomputing science |
The single most important design decision in this project is what is allowed to drive the reconstruction — and methylation is deliberately not.
When you observe two groups of reads with different methylation at a locus, there are at least four possible causes: germline allele-specific methylation, unmasking by loss of heterozygosity, copy-number dosage, and genuine lineage difference. They are not identifiable from the current single-bulk measurement set without orthogonal data or additional assumptions. Using methylation to independently "confirm" a cellular subclone would therefore require an external cellular assignment; conditioning on mutation-defined labels is concordant association, not independent confirmation.
So the backbone is candidate somatic-allele co-occurrence on the same physical molecule. It is non-circular only with respect to methylation-derived labels; it still depends on upstream variant calling, alignment, basecalling and haplotag assumptions. Methylation is retained as a strictly bounded auxiliary signal: it is computed after the genetic candidate structure is fixed, it annotates that structure, and it cannot move a single edge.
Empirically this restraint is justified: of 811 evaluable methylation units, only 3 (0.37 %) supported a robust pattern-conditioned association. This low yield is one reason not to use methylation to choose topology; it is not a test of every possible methylation signal.
Canonical result across 7 datasets (6 biological samples), chr1–22, frozen 2026-08-01. Every layer is auditable and the arithmetic is self-consistent (39,648 + 23,858 = 63,506; + 8,449 = 71,955; + 3,224 + 45 = 75,224; + 10,717 = 85,941;
- 13,014 = 98,955).
| Quantity | Value |
|---|---|
| sSNV dataset records | 469,849 |
k=1 strict read-linkage components |
170,131 / 255,752 strict components (66.5219 %) |
| Analysis units carrying mutations | 85,941 |
| Abandoned at the search-node guard | 10,717 (12.47 %) |
| Units rankable by read-AF | 71,955 |
| Rankable candidate units assigned one rooted-unlabeled candidate-shape signature by the frozen model | 63,506 / 71,955 rankable candidate units (88.2579 %) |
| Confirmed cellular subclones | 0 — see below |
| Confirmed linear ancestry relationships | 0 — model-selected candidate shapes are not biological ancestry truth |
How to read the 88.2579 %. It means: of the 71,955 units that were already rankable, the frozen recurrence-allowed model assigns 63,506 (88.2579 %) one mathematical candidate-shape signature. It is a model-conditional graph statistic, not "88 % of the tumour's evolutionary history is solved". Separately, 170,131 / 255,752 strict components (66.5219 %) are
k=1. The component census cannot be used to derive an isolated-dataset-record percentage.
This project ships a machine-readable claim boundary. The frozen
canonical/cohort_receipt.json records technical_all_pass = true, while
summary/all7_summary.json records validation_evidence_eligible = false
(these are not top-level fields of authority_manifest.json). All named checks recorded by
the frozen cohort receipts passed for those cited frozen artifacts; this is neither current-source
certification nor a production/release-gate PASS, and the batch remains not yet usable as
validation evidence.
| Permitted claims | Explicitly forbidden |
|---|---|
|
|
Within the current single-bulk observation/model, and without integrated CN/LOH, cellular
subclones and linear ancestry cannot be confirmed. Additional independent evidence—such as
single-cell, multi-region, orthogonal copy-number, or purity measurements—may improve
identifiability; no one assay is asserted as the only route. Inferred internal nodes are
labelled inferred rather than asserted.
The table is version-scoped. "Verified" means the listed artifact was run or inspected at the stated audit version; it is not a blanket guarantee for every public command or file.
| Component | Status | Notes |
|---|---|---|
inter_sub_mod |
73afaeac-dirty built and ran with byte-equivalent C++/CMake inputs to release baseline ddd8909a; this is not clean-checkout release certification |
|
| C++ test suite | The locked audit had 0 failures, but release acceptance still requires an exact-commit, repo-external clean build; counts must be generated dynamically | |
| HP/PS sidecars for 7 technical datasets / 6 biological IDs | ✅ complete | 7/7 technical datasets PASS |
| Research exact-PS topology solver | ✅ runnable | Separate research solver path producing the exact-PS funnel above; not inter_sub_mod |
longlineage preflight |
✅ runnable | Validates the 8-role manifest |
longlineage dataset-gate |
In candidate b9aaa12's frozen HCC1395 receipt, the only observed interface producing research artifacts; hardcoded to one dataset, not production science validation |
|
longlineage run / probe |
🔴 blocked | KernelBlocked by design |
longlineage topology output |
🔴 0 candidate units in one frozen receipt | Candidate b9aaa12, HCC1395 dataset-gate scope only; see caveat |
| Writing a tagged BAM | ✅ runnable, still non-production | inter_sub_mod does not write BAM. LongLineage public main 583e03e builds longlineage-tag-bam (CMakeLists.txt:221), merged 2026-08-17 — the earlier baseline snapshot 5daf50f did not. Availability is not a release claim: LongLineage's own gate still reports FAIL, see below |
| Single script spanning both engines | LongLineage ships scripts/run_sample.sh (partition → tagged BAM). No single script yet spans LongLineage and inter_sub_mod; the run order is documented under Running this together with LongLineage |
LongLineage became a public repository on 2026-08-17, and its main was consolidated the
same day: the three agent/public-preview-* branches plus the tag-bam/solver commits were
merged, giving main = 583e03e with SBOM.spdx.json, THIRD_PARTY_NOTICES.md and
docs/release/PUBLIC_SAFETY_RECEIPT.json present. That tree was rebuilt and re-tested in an
isolated worktree: build exit 0, ctest 49/49.
Public and buildable is not the same as released. LongLineage's own
scripts/ci/check_public_preview_gate.sh HEAD still reports FAIL with five open blockers:
license review pending, 4 NO_GO source rows, 21 unapproved source-license rows,
11 NOASSERTION dependencies, and 4 historical blobs carrying developer absolute paths
(its current working tree is clean; the paths survive only in history). P3/P4/P5/P7/P8
remain BLOCKED, production run is intentionally fail-closed, and no production tag or
GitHub Release exists. Read it, build it, run its tests — do not redistribute it as a
license-cleared release.
Historical version notes: b9aaa12 was the public-preview candidate and 5daf50f the earlier
baseline/main snapshot; statements below that are scoped to those commits stay scoped to them.
Its contracts schema-lock artifacts with SHA-256 receipts, and topology_unit records named
abstention states; those contracts do not establish biological correctness.
For candidate b9aaa12's frozen HCC1395 dataset-gate receipt it emits 0 candidate topology units; this does not
generalise to every possible LongLineage run:
Its pipeline gates topology construction on methylation clustering stability, so 66,836 of 79,687 sites (83.9 %) never produce a co-occurrence row at all. This is a direct conflict with the methodological position stated above, and any cross-engine comparison must say so explicitly.
"BLOCKED" means the named P3/P4/P5/P7/P8 gates have not passed. Those gates cover
parity/validation, source-origin, license, dependencies, and release safety. Parts of the
M1/M2/topology kernels were exercised by one named dataset-gate receipt; that does not
establish that every core path, the production entry point, or public redistribution is ready.
Full walkthrough with expected output at every step: How to run
# 0. Clone and check out the immutable commit recorded by the handoff release manifest.
git clone https://github.com/liaoyoyo/InterSubMod.git
cd InterSubMod
HANDOFF_COMMIT="<IMMUTABLE_HANDOFF_COMMIT_SHA>"
git checkout --detach "$HANDOFF_COMMIT"
test "$(git rev-parse HEAD)" = "$HANDOFF_COMMIT"
test -z "$(git status --porcelain)"
# 1. Clean Release build outside the repository; build outputs are not versioned.
REPO_ROOT="$(pwd -P)"
BUILD_ROOT="$(mktemp -d "${TMPDIR:-/tmp}/ism-build.XXXXXXXX")"
cmake -S "$REPO_ROOT" -B "$BUILD_ROOT" -DCMAKE_BUILD_TYPE=Release
cmake --build "$BUILD_ROOT" -j$(nproc)
test -z "$(git -C "$REPO_ROOT" status --porcelain)"
# 2. Verify the build -> derive test/suite counts from this run; require exit 0 and 0 failures
"$BUILD_ROOT/bin/run_tests"
ctest --test-dir "$BUILD_ROOT" --output-on-failure
# 3. Hash-locked Python 3.10 acceptance environment
PYTHON_ENV="$(mktemp -d "${TMPDIR:-/tmp}/ism-python.XXXXXXXX")/venv"
python3.10 -m venv "$PYTHON_ENV"
"$PYTHON_ENV/bin/python" -m pip install --require-hashes \
--requirement "$REPO_ROOT/requirements-ci.lock"
export PATH="$PYTHON_ENV/bin:$PATH"
# 4. Run the tracked tiny synthetic fixture (DEMO/SMOKE only; not science validation).
"$REPO_ROOT/scripts/handoff/run_tiny_public_e2e.sh" --repo-root "$REPO_ROOT" --jobs 4
# 5. Real analysis uses an untracked machine site profile; absolute paths live there/registry only.
SITE_PROFILE="<SITE_PROFILE>"
"$REPO_ROOT/scripts/site/doctor" --profile "$SITE_PROFILE" --mode real-preflight
eval "$("$REPO_ROOT/scripts/site/site_profile.py" shell --profile "$SITE_PROFILE" --sample HCC1395)"
"$BUILD_ROOT/bin/inter_sub_mod" \
--tumor-bam "$TUMOR_BAM" \
--reference "$REFERENCE" \
--vcf "$SOMATIC_VCF" \
--output-dir "${TMPDIR:-/tmp}/ism-real-smoke"The tiny fixture contains tracked synthetic FASTA/VCF/SAM sources and builds BAM/index files in a new temporary directory. A passing receipt proves clone→build→run→schema plumbing only; it is DEMO scope and cannot support biological, benchmark, or release-science claims.
The two engines build independently — neither CMake file references the other — but they compose at run time in one direction:
upstream (neither repo): dorado (MM/ML) → align → LongPhase haplotag → somatic VCF
↓
LongLineage scripts/run_sample.sh → writes BAM aux tags lc:Z lu:Z lv:Z ls:A
↓
InterSubMod inter_sub_mod --tumor-bam <tagged.bam> ...
InterSubMod runs perfectly well without LongLineage. If the BAM never went through
longlineage-tag-bam, the lineage axes are simply skipped rather than collapsed into a single
group — see the lineage-axis comment in include/core/Stats.hpp. The reverse direction,
LongLineage reading an InterSubMod manifest in its src/compat/regional_crosswalk.cpp
(schema_name = intersubmod.layered_sample_output_manifest), is a validation crosswalk and
not a dependency, so do not assume InterSubMod has to run first.
Companion repository, with its own run order and honest gate status:
LongLineage. To check that the LongLineage
commit quoted in this README still matches its public main:
python3 scripts/handoff/check_companion_version.pyThree things that will bite you (all verified against source)
--threadsdefaults to 16, not 1. The help text saysDefault: 1; the actual default inConfig.hppis16. Resource estimates will be off by 16×.- The no-argument
--distance-metricdefault isNHD. The parser replaces theConfig.hppinitializer with its CLI list: explicitNHD,BERNOULLI, or repeated selections are honored. When several metrics are requested, only the first drives clustering; later metrics produce additional distance matrices, so order matters. tree.nwkleaves are reads, not clones. It is a hierarchical clustering tree over read methylation similarity — not a subclonal phylogeny. This is the single most commonly misread output.
Two more version-sensitive details: methylation.csv's first column is a row index rather
than a read name (the binding to reads is positional, with no key check). Frozen release
baseline ddd8909a has a 199-column significance_summary.csv source header, including
VerificationSchemaVersion=2 and RegionStratificationSchemaVersion=1. These are component
schema fields, not a single whole-file layout version; always index by column name and pin
the producing commit.
The Explanation Center (docs/explain/) is the entry point — 17 self-contained pages,
no external dependencies, openable offline.
| Page | For | |
|---|---|---|
| 🗺️ | System overview | Start here — architecture, honest status, funnel, capability boundary |
| ⚙️ | InterSubMod I/O | 3 inputs, 8 stages, 17 output files with real headers, 9 pitfalls |
| 🧬 | LongLineage I/O | Subcommand status, M1→M2→topology chain, artefact contracts |
| 🔬 | Upstream & data | Dorado / ClairS / LongPhase-S / SAVANA, sidecar format, 7 technical datasets / 6 biological IDs |
| 📊 | Analysis & presentation | Which scripts to use, refuse-on-missing HTML generator |
| How to run | Six steps, each with an acceptance check |
Two ways to read the same material. The table above links to the Wiki, which GitHub renders natively and is the quickest to skim. The same content is also served as fully self-contained HTML at liaoyoyo.github.io/InterSubMod — richer typography, interactive fold-out sections, and 37 inline SVG elements in the locked 2026-08-12 deploy (counted as DOM
<svg>elements).docs/explainis the editorial upstream; the Wiki is a manually synchronized derivative and can drift. Publication is a separate step, so compare the deployed commit before assuming byte-level identity.
Pages 01–10 cover the scientific method itself (glossary, ISM core, methylation read/filter, worked case studies, statistical division of labour, capability vs. narrative).
Two patterns in this repository generalise beyond genomics.
1 · Streaming as a supported workflow mode. The frozen nine-field sidecars occupy
5.83 GiB and retain 100% of the fields required by that downstream contract. FIFO mode can
avoid materialising a new tagged BAM for that run; it does not mean no tagged BAM has ever
existed. The 2026-08-13 inventory identifies seven historical PRE-FIX paired_full BAMs
totaling 1,840,983,466,353 bytes (1.67436 TiB), plus seven paired_pileup BAMs; all 14 total
3,709,322,840,333 bytes. Exact paths/bytes/mtimes and a sampled-content set identity are known,
but full-file hashes and generation-level correspondence are not, so all remain
PARTIAL/NON_FINAL. Dividing the seven paired_full stat sizes by the seven current sidecar
stat sizes gives 294.2669×; this is only a cross-generation storage-footprint quotient, not a
causal compression result, a byte-equivalent replacement claim, or frozen authority. The old
287× figure is incorrect.
2 · Refuse missing required metrics. The audited workstation generator takes a declarative spec and refuses to render (exit 3) when a required metric is missing. This fail-closed behaviour covers that generator and its declared fields; it is not proof that every public number or document is automatically validated.
- C++17 compiler, CMake ≥ 3.14, OpenMP
- htslib (BAM/VCF I/O), Eigen
- samtools for building/indexing the tracked tiny synthetic fixture
- Python 3.10 for acceptance, installed from
requirements-ci.lockwith--require-hashes;requirements.txtis the input/general-plotting dependency list, not the handoff acceptance environment - Reference FASTA with
.fai(a missing index is a hard error) - Input BAM must carry
MM/MLmethylation tags (reads without them are dropped)
include/core/ src/core/ C++17 analysis engine
scripts/ Python analysis entry points
docs/explain/ Explanation Center (17 standalone pages)
docs/images/ Figures used by this README and the wiki
tools/ Rendering, QA and extraction utilities
state/ Research cycle state machine
This README is bounded by the
2026-08-12 public-claim audit.
The frozen exact-PS counts were checked against the authority manifest and denominator
registry; the tracked C++ core was freshly built and tested. The tracked tiny synthetic fixture
is shipped for DEMO/SMOKE only, while real tumour BAM/reference/VCF inputs are not; GitHub About/Wiki/Pages have independent publication state, and LongLineage
capabilities are pinned to the commits named above. Therefore no blanket "every number and
command was verified" claim is made. Figures are extracted from docs/explain/ by
tools/extract_svg_for_github.py; regenerate them after editing their source pages.
| Artifact / claim family | Verification identity and date | Reproducible check and result | Scope and known failure |
|---|---|---|---|
| Frozen exact-PS funnel | Frozen authority artifacts, rechecked 2026-08-12 | Manifest/hash census plus independent denominator recomputation; exact counts reproduced | Frozen 7-dataset analysis only; it is not the inter_sub_mod CLI and does not identify cellular clones |
| Frozen release-baseline source | ddd8909a |
The tracked source header has 199 columns; current test/suite counts must be generated dynamically from an exact-commit clean run | Historical 73afaeac-dirty audit inputs were byte-equivalent for C++/CMake and its 0-failure receipt is historical evidence only, not clean-checkout release certification |
| Public quick start | Current source, reviewed 2026-08-13 | Build/test plus the tracked tiny synthetic fixture reproduce clone→build→run→schema plumbing | The fixture is DEMO only; no public real tumour BAM/reference/VCF is shipped, so no biological result is claimed |
| LongLineage status | Private repo; public-preview candidate b9aaa12; private baseline snapshot 5daf50f; frozen HCC1395 evidence |
Source/commit comparison plus one-dataset frozen artifact audit | NOT_READY; P3/P4/P5/P7/P8 BLOCKED; tagged-BAM belongs only to the candidate; no seven-technical-dataset/six-biological-ID runtime, memory, or topology validation |
| GitHub surfaces | Main-repo source corrected locally 2026-08-13 | P0 claim registry and source checks in this correction cycle | About, Wiki, Pages and default-branch publication are separate external actions; source edits alone do not prove deployment |
The exact commands and captured output are preserved in the command receipts; the per-claim correction status is tracked in the P0 correction cycle.
Known gaps, stated honestly: CN/LOH are not integrated into the frozen candidate reconstruction and its inferences are not CN/LOH-corrected; InterSubMod's optional LOH BED is annotation/ stratification only. LongLineage's seven-technical-dataset/six-biological-ID runtime and memory ceiling have never been measured and its own documentation forbids extrapolating from one sample.






