Raw paired-end metagenomic reads were processed on CSC Puhti before taxonomic and functional profiling. Scripts are in raw_reads/.
Pipeline: nf-core/taxprofiler v1.1.5 (Nextflow).
Steps enabled in raw_reads/taxprofiler_runs.sh:
- QC — FastQC
- Trimming — fastp (adapters and low-quality bases)
- Complexity filter — BBduk
- Host removal — Bowtie2 against human reference T2T-CHM13v2
- Taxonomy — MetaPhlAn4 (
merge_metaphlan_tables.pyfor merged profiles)
Run example: sbatch raw_reads/taxprofiler_runs.sh (edit samplesheet and config paths for your Puhti project).
More setup notes: workflow_metagenome.
HUMAnN3 via Snakemake (raw_reads/Snakefile_humann.txt, submit with raw_reads/run_humann_workflow.sh) on host-depleted FASTQs (*.unmapped_1.fastq.gz). Post-processing: raw_reads/humann_modif.sh.
MetaPhlAn abundance tables used in the Quarto pipeline are cleaned with:
remove_columns_metaphlan_db_meta4_combined_reports.R— drop sample columns and writelatest_metaphlan_db_meta4_combined_reports.txt
Read deduplication was not included in this workflow.