Skip to content

Commit 985b4c3

Browse files
committed
add rna usecase
1 parent 8b13f79 commit 985b4c3

3 files changed

Lines changed: 5082 additions & 45 deletions

File tree

‎doc/_static/CGGA_P438_cancer_report.html‎

Lines changed: 4965 additions & 0 deletions
Large diffs are not rendered by default.

‎doc/usecase/case_one.md‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -49,8 +49,8 @@ Next, create a CSV file named pipe_wes.csv in the ~/projects/CGGA_WES directory
4949

5050
```
5151
Tumor_R1_file_path,Tumor_R2_file_path,Normal_R1_file_path,Normal_R2_file_path,Sample_name,Target_file_bed,Project
52-
T_CGGA_D14_r1.fq.gz,T_CGGA_D14_r2.fq.gz,B_CGGA_D14_r1.fq.gz,B_CGGA_D14_r1.fq.gz,CGGA_D14,target.bed,CGGA_WES
53-
T_CGGA_653_r1.fq.gz,T_CGGA_653_r2.fq.gz,B_CGGA_653_r1.fq.gz,B_CGGA_653_r1.fq.gz,CGGA_653,target.bed,CGGA_WES
52+
~/projects/CGGA_WES/data/T_CGGA_D14_r1.fq.gz,~/projects/CGGA_WES/data/T_CGGA_D14_r2.fq.gz,~/projects/CGGA_WES/data/B_CGGA_D14_r1.fq.gz,~/projects/CGGA_WES/data/B_CGGA_D14_r1.fq.gz,CGGA_D14,target.bed,CGGA_WES
53+
~/projects/CGGA_WES/data/T_CGGA_653_r1.fq.gz,~/projects/CGGA_WES/data/T_CGGA_653_r2.fq.gz,~/projects/CGGA_WES/data/B_CGGA_653_r1.fq.gz,~/projects/CGGA_WES/data/B_CGGA_653_r1.fq.gz,CGGA_653,target.bed,CGGA_WES
5454
```
5555

5656
## Write an Snakemake file from template
@@ -169,4 +169,10 @@ nohup snakemake --profile workflow/config_slurm \
169169
-j 30 --printshellcmds -s snake_wes.smk --use-singularity \
170170
--singularity-args "--bind /public/home/:/public/home/,/public/ClinicalExam:/public/ClinicalExam" \
171171
--latency-wait 300 --use-conda >> wes.log
172-
```
172+
```
173+
174+
### Output
175+
176+
### case report
177+
There is a example case report of CGGA_P438
178+
<a href="../_static/CGGA_P438_cancer_report.html">example report HTML</a>

‎doc/usecase/case_two.md‎

Lines changed: 108 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,67 @@
11
# Use case II: Fusion genes detection from multiple myeloma patient RNA-seq
2-
32
## Background
3+
​​Clinical Applications of RNA-Seq in Diagnostic Testing​​
4+
5+
RNA sequencing (RNA-Seq) is a high-throughput transcriptome profiling technology that enables comprehensive analysis of gene expression, splicing variants, fusion events, and novel transcripts. In clinical diagnostics, it serves as a powerful tool for:
6+
7+
∙
8+
​​Cancer Subtyping​​: Identifying tumor-specific gene expression signatures, fusion genes (e.g., BCR-ABL1), and aberrant splicing events to guide targeted therapies.
9+
10+
∙
11+
​​Rare Disease Diagnosis​​: Detecting dysregulated pathways and aberrant expression in Mendelian disorders where DNA-based tests are inconclusive.
12+
13+
∙
14+
​​Infectious Disease Characterization​​: Profiling host-pathogen interactions and pathogen expression in complex infections.
15+
16+
∙
17+
​​Biomarker Discovery​​: Validating expression-based biomarkers for disease monitoring and treatment response.
18+
19+
20+
## Setup a project folder
21+
````{note}
22+
Before starting the analysis, please ensure that you have set up the analysis environment using the build_conda_env.sh script.
23+
````
24+
25+
Create a folder named project/CGGA_WES in your home directory and activate the Clindet conda environment.
26+
27+
```{code} bash
28+
mkdir -p ~/projects/MM_RNA
29+
cd ~/projects/MM_RNA
30+
conda activate clindet
31+
```
32+
## Download data and
33+
34+
Download Multiple myeloma and COLO829 cellline RNA-seq data from the SRA database using wget and prepare the sample information file, make sure fastq-dump are in in $PATH (if don't install it first)
35+
36+
```{code} bash
37+
cd ~/projects/MM_RNA
38+
mkdir -p data && cd data
39+
## Methods one multiple myeloma RNA-seq data
40+
wget -q -c -O A26.11 https://sra-pub-run-odp.s3.amazonaws.com/sra/SRR12099713/SRR12099713
41+
wget -q -c -O A27.19 https://sra-pub-run-odp.s3.amazonaws.com/sra/SRR12099714/SRR12099714
42+
wget -q -c -O A28.15 https://sra-pub-run-odp.s3.amazonaws.com/sra/SRR12099715/SRR12099715
43+
44+
fastq-dump --gzip -O /public/ClinicalExam/lj_sih/projects/project_clindet/data/GSE153380 --split-3 ./A26.11
45+
fastq-dump --gzip -O /public/ClinicalExam/lj_sih/projects/project_clindet/data/GSE153380 --split-3 ./A27.19
46+
fastq-dump --gzip -O /public/ClinicalExam/lj_sih/projects/project_clindet/data/GSE153380 --split-3 ./A28.15
47+
48+
```
49+
50+
Next, create a CSV file named pipe_rna.csv in the ~/projects/MM_RNA directory with the following content:
51+
52+
```
53+
Tumor_R1_file_path,Tumor_R2_file_path,Normal_R1_file_path,Normal_R2_file_path,Sample_name,Target_file_bed,Project
54+
~/projects/MM_RNA/data/A26.11_1.fastq.gz,~/projects/MM_RNA/data/A26.11_2.fastq.gz,MF1
55+
~/projects/MM_RNA/data/A27.19_1.fastq.gz,~/projects/MM_RNA/data/A27.19_2.fastq.gz,MS3
56+
~/projects/MM_RNA/data/A28.15_1.fastq.gz,~/projects/MM_RNA/data/A28.15_2.fastq.gz,CD1
57+
```
58+
59+
## Write an Snakemake file from template
60+
For this project, modify the sample sheet and create a new Snakemake file named **snake_rna.smk** (see below). Set the following parameters in the Snakemake file:
61+
62+
1. **configfile (str)**: config file for softwares and resource parameters.
63+
1. **stage (list)**: analysis steps. avaiable options:`['RSEM','arriba','TRUST4','samlom','kallisto']`
464

5-
## Download data
665

766
## write Snakemake file
867
For this project, we need change the sample sheet info.
@@ -12,67 +71,74 @@ For this project, we need change the sample sheet info.
1271
```{code} python
1372
1473
import pandas as pd
15-
samples_info = pd.read_csv('',index_col='Sample_name') # set sample sheet path
16-
17-
unpaired_samples = samples_info.loc[pd.isna(samples_info['Normal_R1_file_path'])].index.tolist()
18-
paired_samples = samples_info.loc[~pd.isna(samples_info['Normal_R1_file_path'])].index.tolist()
74+
samples_info = pd.read_csv('./pipe_rna.csv',index_col='Sample_name')
75+
unpaired_samples = samples_info.loc[pd.isna(samples_info['R2_file_path'])].index.tolist()
76+
paired_samples = samples_info.loc[~pd.isna(samples_info['R1_file_path'])].index.tolist()
1977
20-
configfile: "" # set config file path
21-
22-
project = samples_info["Project"].unique().tolist()[0]
23-
genome_version = 'WBcel235' # set genome version
78+
configfile: "/public/ClinicalExam/lj_sih/projects/project_clindet/build_log/config.yaml"
2479
2580
import os
2681
if not os.path.exists("logs/slurm"):
2782
os.makedirs("logs/slurm")
2883
29-
pre_pon_db = False
30-
31-
if not os.path.exists('analysis/pindel_normal/log'):
32-
os.makedirs('analysis/pindel_normal/log')
33-
3484
groups = ['NC','T']
35-
36-
germ_caller_list = ['caveman']
37-
caller_list = ['strelkasomaticmanta','caveman','muse','cgppindel_filter']
38-
39-
recall_pon = False
40-
recall_pon_pindel = False
41-
42-
recal = False
85+
stages = ['RSEM','arriba','TRUST4','samlom','kallisto']
86+
caller_list = ['sentieon_anno_rnaedit','Mutect2_filter']
87+
project = 'RNA'
88+
genome_version = 'b37'
89+
90+
rna_res_list = [
91+
##### for isoform expression ######
92+
"{project}/{genome_version}/results/summary/RSEM/{sample}/{sample}.genes.results" if 'RSEM' in stages else None,
93+
##### ka
94+
"{project}/{genome_version}/results/summary/kallisto/{sample}/abundance.tsv" if 'kallisto' in stages else None,
95+
##### for fusion gene detection #####
96+
"{project}/{genome_version}/results/fusion/{sample}_arriba_fusion.tsv" if 'arriba' in stages else None,
97+
##### for TRUST4 immu analysis #####
98+
"{project}/{genome_version}/results/IG/TRUST4/{sample}_report.tsv" if 'TRUST4' in stages else None,
99+
#### Case report #####
100+
]
101+
rna_res_list = list(filter(None, rna_res_list))
43102
rule all:
44103
input:
45104
## paired sample
46-
expand([
47-
"{project}/{genome_version}/results/dedup/paired/{sample}-{group}.sorted.bam",
48-
"{project}/{genome_version}/results/sv/paired/DELLY/{sample}/SV_delly_{sample}_filter.vcf",
49-
"{project}/{genome_version}/results/vcf/paired/{sample}/strelkasomaticmanta.vcf",
50-
"{project}/{genome_version}/results/vcf/paired/{sample}/muse.vcf",
51-
'{project}/{genome_version}/logs/paired/caveman_{sample}.log',
52-
],
105+
expand(rna_res_list,
106+
# sample = paired_samples,
107+
sample = ['CD1','COLO829'],
53108
project = project,
54-
genome_version = genome_version,
55-
sample = paired_samples,
56-
group = groups,
57-
caller = caller_list)
58-
59-
include: "workflow/WGS/Snakefile" # the relative path of clindet workflow WGS subfolder snakefile
109+
genome_version = genome_version
110+
),
111+
112+
##### Modules #####
113+
include: "workflow/RNA/Snakefile"
60114
61115
62116
```
63117
:::
64118

65119
## Run clindet
120+
There is two way you can run clindet
121+
1. run on a local server
122+
2. submit to HPC through slurm
123+
124+
### Run on local node
125+
```{code} bash
126+
nohup snakemake -j 30 --printshellcmds -s snake_rna.smk \
127+
--use-singularity --singularity-args "--bind /your/home/path:/your/home/path" \
128+
--latency-wait 300 --use-conda >> rna.log
129+
```
66130

67-
``` bash
131+
### Submit to HPC use slurm
132+
we provide a slurm config.yaml under clindet/workflow/config_slurm folder.
133+
```{code} bash
68134
nohup snakemake --profile workflow/config_slurm \
69-
-j 30 --printshellcmds -s snake_wgs_worm.smk \
70-
--use-singularity \
71-
--singularity-args "--bind /public/home/:/public/home/,/public/ClinicalExam:/public/ClinicalExam" \
72-
--latency-wait 300 --use-conda -
73-
135+
-j 30 --printshellcmds -s snake_rna.smk --use-singularity \
136+
--singularity-args "--bind /your/home/path:/your/home/path" \
137+
--latency-wait 300 --use-conda >> rna.log
74138
```
75139

140+
### Output
141+
76142
## Results
77143

78144
```{image} ../img/usecase/usecase_two/Fusion_gene.jpeg

0 commit comments

Comments
 (0)