Skip to content

Commit 8b13f79

Browse files
committed
update usecase
1 parent 1c559bf commit 8b13f79

7 files changed

Lines changed: 163 additions & 252 deletions

File tree

‎doc/Quick/index.rst‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,5 @@ Quick Start
77

88
intro
99
install.md
10-
config.md
1110
use_case
12-
workflow
11+
workflow.md

‎doc/Quick/install.md‎

Lines changed: 7 additions & 177 deletions
Original file line numberDiff line numberDiff line change
@@ -4,190 +4,20 @@
44

55
## install conda, Docker and SingularityCE
66

7-
To build the complex clindet analysis environment, you need install
7+
To build the complex Clindet analysis environment, you need to install Conda, Docker, and SingularityCE.
88

99
## Install Clindet
10-
1110
### clone clindet from github
1211

1312
```bash
14-
git clone clindet.git
13+
git clone https://github.com/zyllifeworld/clindet.git
1514
cd clindet
1615
```
17-
## Download and config Clindet pre-built singularity Image
18-
### Download from zendo
19-
20-
1. cgpindel
21-
2. caveman
22-
3. brass
23-
4. freec
24-
5. cgpwgs
25-
6. ascat
26-
7. arriba
27-
8. lofreq
28-
9. muse
29-
10. gridss
30-
11. hmftools
31-
12. conpair
32-
33-
34-
### config image path in config.yaml
35-
36-
37-
## Download and config Genome refernce file (eg, human b37)
38-
39-
clindet 参考GATK最佳实践的方法对测序fastq文件进行预处理。使用者可以从GATK的网站下载所需的文件,具体可参照[GATK Resource bundle](https://gatk.broadinstitute.org/hc/en-us/articles/360035890811-Resource-bundle)。包括fasta,GTF文件等。在clindet中,我们提供了人类基因组版本b37各文件的下载脚本用于自动下载,具体步骤如下:
40-
41-
### 软件安装
42-
43-
* **conda 环境 gsutils安装与配置**
44-
45-
```bash
46-
conda env create -f env/gsutils.yaml
47-
conda activate gsutils
48-
```
49-
50-
* **conda 环境 clindet 安装与配置**
51-
52-
```bash
53-
conda env create -f env/clindet.yaml
54-
conda activate clindet
55-
```
56-
57-
* **下载并安装GATK**
58-
59-
参照[GATK官网](https://github.com/broadinstitute/gatk)的说明进行GATK Toolkit的安装与配置,建议将该软件安装到家目录下的softwares文件夹中
60-
61-
* **下载并安装picard**
62-
63-
参照[picard官网](https://broadinstitute.github.io/picard/)进行picard软件的安装与配置
64-
65-
### download B37
66-
67-
On folder clindet, make a folder to store genome fastq and other files for b37
68-
69-
```bash
70-
mkdir -p reference/human/b37
71-
cd reference/human/b37
72-
```
73-
74-
using following scripts to download files
75-
76-
```bash
77-
gsutil -m cp -r \
78-
"gs://gcp-public-data--broad-references/hg19/v0/1000G_omni2.5.b37.vcf.gz" \
79-
"gs://gcp-public-data--broad-references/hg19/v0/1000G_omni2.5.b37.vcf.gz.tbi" \
80-
"gs://gcp-public-data--broad-references/hg19/v0/1000G_phase1.snps.high_confidence.b37.vcf.gz" \
81-
"gs://gcp-public-data--broad-references/hg19/v0/1000G_phase1.snps.high_confidence.b37.vcf.gz.tbi" \
82-
"gs://gcp-public-data--broad-references/hg19/v0/1000G_reference_panel" \
83-
"gs://gcp-public-data--broad-references/hg19/v0/Axiom_Exome_Plus.genotypes.all_populations.poly.vcf.gz" \
84-
"gs://gcp-public-data--broad-references/hg19/v0/Axiom_Exome_Plus.genotypes.all_populations.poly.vcf.gz.tbi" \
85-
"gs://gcp-public-data--broad-references/hg19/v0/ExomeContam.vcf" \
86-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.cdna.all.fa" \
87-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.cds.all.fa" \
88-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.cloud_references.json" \
89-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.contam.UD" \
90-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.contam.V" \
91-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.contam.bed" \
92-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.contam.mu" \
93-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.dbsnp138.vcf" \
94-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.dbsnp138.vcf.idx" \
95-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard0.tsv" \
96-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard1.tsv" \
97-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard10.tsv" \
98-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard11.tsv" \
99-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard2.tsv" \
100-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard3.tsv" \
101-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard4.tsv" \
102-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard5.tsv" \
103-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard6.tsv" \
104-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard7.tsv" \
105-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard8.tsv" \
106-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.delly_exclusionRegions.shard9.tsv" \
107-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.dict" \
108-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta" \
109-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.64.amb" \
110-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.64.ann" \
111-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.64.bwt" \
112-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.64.pac" \
113-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.64.sa" \
114-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.alt" \
115-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.amb" \
116-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.ann" \
117-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.bwt" \
118-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.fai" \
119-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.pac" \
120-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.rbwt" \
121-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.rpac" \
122-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.rsa" \
123-
"gs://gcp-public-data--broad-references/hg19/v0/Homo_sapiens_assembly19.fasta.sa" \
124-
.
125-
126-
gsutil -m cp \
127-
"gs://gatk-best-practices/somatic-b37/Mutect2-WGS-panel-b37.vcf" \
128-
"gs://gatk-best-practices/somatic-b37/Mutect2-WGS-panel-b37.vcf.idx" \
129-
"gs://gatk-best-practices/somatic-b37/Mutect2-exome-panel.vcf" \
130-
"gs://gatk-best-practices/somatic-b37/Mutect2-exome-panel.vcf.idx" \
131-
.
132-
133-
134-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/Mills_and_1000G_gold_standard.indels.b37.vcf.gz
135-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/Mills_and_1000G_gold_standard.indels.b37.vcf.gz.md5
136-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase1.indels.b37.vcf.gz
137-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase1.indels.b37.vcf.gz.md5
138-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase1.snps.high_confidence.b37.vcf.gz
139-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase1.snps.high_confidence.b37.vcf.gz.md5
140-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/hapmap_3.3.b37.vcf.gz
141-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/hapmap_3.3.b37.vcf.gz.md5
142-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_omni2.5.b37.vcf.gz
143-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_omni2.5.b37.vcf.gz.md5
144-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase3_v4_20130502.sites.vcf.gz
145-
wget -c ftp://gsapubftp-anonymous@ftp.broadinstitute.org/bundle/b37/1000G_phase3_v4_20130502.sites.vcf.gz.md5
146-
147-
```
148-
149-
## Resource bundle
150-
151-
## edit config file
152-
153-
config文件主要用来包括三大部分参数的配置,可以在包括
154-
155-
* Resource
156-
用来记录各版本参考基因组的fasta, GTF,Panel of Normal文件的绝对路径信息
157-
* softwares
158-
用来记录分析软件的绝对路径,conda环境,参数等信息
159-
* singularity
160-
用来记录singularity封装容器的绝对路径与运行参数
161-
162-
以b37版本为例,在使用下载[下载脚本](#download B37)下载参考基因组后(也可以使用其他方法自行从GATK中下载),首先需要在workflow/config文件夹中创建config.yaml文件(以该文件夹下的config_local_test.yaml作为模版),在resource section进行如下配置:
163-
164-
```yaml
165-
resources:
166-
b37:
167-
REFFA: "/Your_file_path/reference/b37/Homo_sapiens_assembly19.fasta"
168-
GENOME_BED: ""
169-
GTF: ""
170-
WES_PON: "/Your_file_path/reference/b37/Mutect2-exome-panel.vcf"
171-
WES_BED: ""
172-
WGS_PON: "/Your_file_path/reference/b37/Mutect2-WGS-panel-b37.vcf"
173-
DBSNP: "/Your_file_path/reference/b37/Homo_sapiens_assembly19.dbsnp138.vcf"
174-
DBSNP_GZ: "/Your_file_path/reference/b37/Homo_sapiens_assembly19.dbsnp138.vcf.gz"
175-
MUTECT2_VCF: "/Your_file_path/af-only-gnomad.raw.sites.b37.vcf.gz"
176-
REFFA_DICT: "/Your_file_path/reference/b37/Homo_sapiens_assembly19.fasta"
177-
MUTECT2_germline_vcf: "/Your_file_path/af-only-gnomad.raw.sites.b37.vcf.gz"
178-
```
179-
180-
需将以上字段自行配置为下载后软件的位置
181-
182-
其中GTF可自行从[GeneCode](https://www.gencodegenes.org/human/release_46lift37.html)网站下载,但需要从GTF文件中去掉’chr' prefix。人类各版本基因组的不同可以参考[This Post](https://www.gencodegenes.org/human/release_46lift37.html).
183-
184-
GENOME_BED,与WES_BED主要用来记录染色体与基因外显子区间的坐标文件,可为空。
185-
186-
其余参数的具体说明可参考
187-
16+
### Run build_conda_env.sh to
17+
Clindet provides a bash script to set up the computational environment required for running the software, as well as to download configuration files needed for various tools that use the human b37 reference genome (e.g., VCF, BED files). Run the script build_conda_env.sh to complete this setup.
18818

189-
## VEP setup
190-
Install VEP use conda/ or use VEP
19+
### Modify the config.yaml file
20+
Replace the placeholder '/AbsoPath/of/clindet/folder' in the **clindet/workflow/config/config_local_test.yaml** file located in the Clindet software directory with the absolute path to your Clindet folder, for example: /home/users/softwares/clindet. save the modified file as config.yaml.
19121

192-
conda
19322

23+
You are now ready to start your analysis tasks!

‎doc/Quick/intro.md‎

Lines changed: 11 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -7,15 +7,15 @@
77
[![Github](https://img.shields.io/github/stars/clindet/clindet?style=social)](https://github.com/clindet/clindet/stargazers)
88

99
## Introduction
10-
Clindet is an Next generation High throughout sequencing data analysis wrokflow for clinical applications (eg. DNA-seq )
11-
**Clindet** is a snakemake pipeline for the comprehensive analysis of cancer genomes and transcriptomes using multiple state-of-art softwares to get consensus results. The pipeline
12-
supports a wide range of experimental setups:
10+
Clindet is a next-generation high-throughput sequencing data analysis workflow designed for clinical applications (e.g., DNA-seq, RNA-seq).
11+
**Clindet** is a Snakemake pipeline for comprehensive analysis of cancer genomes and transcriptomes, integrating multiple state-of-the-art tools to generate consensus results. The pipeline supports a wide range of experimental setups, including:
12+
13+
- FASTQ input files
14+
- Whole genome sequencing (WGS), whole transcriptome sequencing (WTS), and targeted/panel sequencing
15+
- Paired tumor/normal and tumor-only sample configurations
16+
- Most GRCh37 and GRCh38 reference genome builds
17+
- Non-human species (e.g., mouse, worm)
1318

14-
- FASTQ
15-
- WGS (whole genome sequencing), WTS (whole transcriptome sequencing), and targeted / panel sequencing
16-
- Paired tumor / normal and tumor-only sample setups
17-
- Most GRCh37 and GRCh38 reference genome builds
18-
- Non-human species(eg, mouse, worm).
1919

2020
## Pipeline overview
2121

@@ -29,7 +29,7 @@ supports a wide range of experimental setups:
2929
## Steps
3030
- Quality Control: ([Conpair](https://github.com/nygenome/Conpair), [fastp](https://github.com/OpenGene/fastp))
3131
- Read alignment: ([BWA-MEM2](https://github.com/bwa-mem2/bwa-mem2) (DNA), [STAR](https://github.com/alexdobin/STAR) (RNA))
32-
- Read post-processing: ([GATK MarkDuplicates](https://gatk.broadinstitute.org/hc/en-us/articles/360037052812-MarkDuplicates-Picard) (DNA,RNA) )
32+
- Read post-processing: ([GATK MarkDuplicates](https://gatk.broadinstitute.org/hc/en-us/articles/360037052812-MarkDuplicates-Picard) , [GATK BaseRecalibrator](https://gatk.broadinstitute.org/hc/en-us/articles/360037052812-MarkDuplicates-Picard) , [GATK ApplyBQSR](https://gatk.broadinstitute.org/hc/en-us/articles/360037052812-MarkDuplicates-Picard))
3333
- SNV, MNV, INDEL calling:
3434
([SAGE](https://github.com/hartwigmedical/hmftools/tree/master/sage), [HaplotypeCaller](https://github.com/broadinstitute/gatk),
3535
[Mutect2](https://github.com/broadinstitute/gatk),
@@ -69,9 +69,6 @@ supports a wide range of experimental setups:
6969
- Oncoviral detection: [VIRUSbreakend](https://github.com/PapenfussLab/gridss)\*, [VirusInterpreter](https://github.com/hartwigmedical/hmftools/tree/master/virus-interpreter)\*
7070
- Telomere characterisation: [TEAL](https://github.com/hartwigmedical/hmftools/tree/master/teal)\*
7171
- Immune analysis: [LILAC](https://github.com/hartwigmedical/hmftools/tree/master/lilac), [CIDER](https://github.com/hartwigmedical/hmftools/tree/master/cider), [NEO](https://github.com/hartwigmedical/hmftools/tree/master/neo)\*
72-
- Mutational signature fitting: [SIGS](https://github.com/hartwigmedical/hmftools/tree/master/sigs)\*
73-
- HRD prediction: [CHORD](https://github.com/hartwigmedical/hmftools/tree/master/chord)\*
74-
- Tissue of origin prediction: [CUPPA](https://github.com/hartwigmedical/hmftools/tree/master/cuppa)\*
7572

7673
- Summary report: [ORANGE](https://github.com/hartwigmedical/hmftools/tree/master/orange)
7774

@@ -84,7 +81,7 @@ supports a wide range of experimental setups:
8481

8582
Create a samplesheet with your inputs (WGS/WES *fastq in this example):
8683

87-
```csv
84+
```{code}
8885
Tumor_R1_file_path,Tumor_R2_file_path,Normal_R1_file_path,Normal_R2_file_path,Sample_name,Target_file_bed,Project
8986
Patient1_T_R1.fq.gz,Patient1_T_R2.fq.gz,Patient1_N_R1.fq.gz,Patient1_N_R2.fq.gz,Patient1,target.bed,WES
9087
Patient2_T_R1.fq.gz,Patient2_T_R2.fq.gz,,,Patient2,target.bed,WES
@@ -114,10 +111,9 @@ the [National Research Center for Translational Medicine at Shanghai](https://gi
114111

115112
We thank the following organisations and people for their extensive assistance in the development of this pipeline,
116113
listed in alphabetical order:
117-
118-
- [Hartwig Medical Foundation Australia](https://www.hartwigmedicalfoundation.nl/en/partnerships/hartwig-medical-foundation-australia/)
119114
- [Broad Institute](https://www.broadinstitute.org/)
120115
- [German Cancer Research Center](https://www.dkfz.de/en/)
116+
- [Hartwig Medical Foundation Australia](https://www.hartwigmedicalfoundation.nl/en/partnerships/hartwig-medical-foundation-australia/)
121117
- [Wellcome Sanger Institute](https://www.sanger.ac.uk/)
122118
- [New York Genome Center](https://www.nygenome.org/)
123119
- JianFeng Li

‎doc/Quick/use_case.md‎

Lines changed: 1 addition & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -1,14 +1,2 @@
11
# Use case
2-
3-
4-
## case: WES
5-
6-
在本案例中我们拟采用人类胶质瘤母细胞的全外显子组配对样本数据进行检测,首先进行数据的下载
7-
8-
### 准备样本的sample_sheet.xlsx文件
9-
10-
### write snakefile for project
11-
12-
由于不同项目的参数配置以及分析过程不同,强烈建议为每一个分析项目编写一个Snakefile,
13-
14-
## case two RNA: fusion
2+
See use case section: <project:../usecase/index.rst> for how to start an analysis task.

‎doc/index.rst‎

Lines changed: 0 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -5,12 +5,6 @@
55
66
Clindet documentation
77
=====================
8-
9-
Add your content using ``reStructuredText`` syntax. See the
10-
`reStructuredText <https://www.sphinx-doc.org/en/master/usage/restructuredtext/index.html>`_
11-
documentation for details.
12-
13-
148
.. toctree::
159
:maxdepth: 2
1610
:caption: Contents:

0 commit comments

Comments
 (0)