This repository contains all the code necessary to reproduce the figures and analyses presented in the scATAcat Manuscript. Below is a brief overview of the repository structure:
-
Preprocessing
- Obtaining ENCODE cCRE coverages of all the datasets
- Doublet detection for single-cell ATAC-seq datasets
- Generating synthetic bulk dataset for fesibility study
- Differential accesibility analysis for the prototype bulk ATAC-seq samples
- Peak calling and gene score calculation via ArchR
-
Notebooks
- Jupyter Notebooks to reproduce figures in the manuscript
- For each dataset, there is seperate folder for the following methods:
- scATAcat
- Marker-based annotation
- Seurat label-transfer
- Cellcano
- EpiAnno
Supplementary tablesincludes notebooks reproducing Supplementary tables
-
Results
- This folder contains all figures and output files generated by the notebooks
-
Data
- Placeholder for the data folder. The content of this file can be access from Zenado repository
If you use the code or data from this repository in your research, please cite our manuscript: Altay, Aybuge, and Martin Vingron. "scATAcat: Cell-type annotation for scATAC-seq data." bioRxiv (2024): 2024-01.