Skip to content

Adopt combined-index Bismark alignment modes / rethink the --multicore default (Bismark 3.1.0) #615

Description

@FelixKrueger

Context

Since the Rust rewrite, the Bismark aligner offers combined-index alignment modes that are generally faster and lighter than the classic per-strand + --multicore model. The Bismark docs now give a per-library-type decision guide: Choosing an alignment mode.

methylseq currently drives alignment with the classic model (the bismark/align module auto-computes --multicore from task.cpus). With Bismark 3.1.0 now in the pipeline (#614), it's worth deciding whether/how methylseq should adopt the newer modes.

The recommendations (summary)

Combined modes are driven by Bowtie 2 threads (-p), and --multicore is rejected in combined mode.

Library ⚡ Fastest / least CPU 🔒 Byte-identical to Perl
Directional --combined_index (one pass, tune -p) — ~22–28% less CPU standard, 2 instances
Non-directional --combined_index_sequential — fastest, leanest RAM (~11 GB), byte-identical to the parallel combined run standard, 4 instances
PBAT --combined_index (one pass, tune -p) standard, 2 instances

What adopting this would involve in methylseq

  • A one-time combined-index build in genome prep (bismark prepare --combined_genome; ~+1.3 GB, extra prep time).
  • Selecting the alignment flag per library type (directional/PBAT vs non-directional) in the alignment subworkflow.
  • Replacing the --multicore resource model with a -p-based one (threads through task.cpus / process labels).

Key decision — and why I'm opening this rather than just PR'ing it

The combined index for directional/PBAT is concordance-gated, not byte-identical to Perl v0.25.1 (≈1 read in 10⁴ placed differently but equally validly); only non-directional-sequential is byte-identical. So this changes default alignment output for the common case. That's a deliberate defaults change for a flagship pipeline, so I'd like to agree the shape first:

  1. Default flip vs opt-in? Make combined-index the default, or expose it as a parameter and keep the faithful (byte-identical) path as default?
  2. Per-library-type auto-selection in the subworkflow — reasonable, or too magic?
  3. Any concern about the extra combined-index build step / disk (sequential's small BGZF spill)?

Happy to implement whichever shape we land on. Follow-up to #614 (the straight 3.0.0→3.1.0 bump, which is intentionally byte-identical and separate from this).

cc @pinin4fjords

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions