Spatial niche identification aims to partition tissue into multicellular microenvironments that are defined by coordinated cell type composition and spatially structured cellular states, using spatially resolved expression measurements. Across tissues, such microenvironments can reflect recurrent cellular neighborhoods, context-dependent state programs, and local cell–cell interactions that are not fully captured by gene expression alone or by coarse anatomical landmarks. Niches may align with anatomy in some settings, but they can also be sharply separated yet internally heterogeneous, appear as non-contiguous islands embedded within larger compartments, or vary continuously along gradients. These properties make it difficult to infer method performance from a single reference setting, motivating a benchmark that jointly probes multiple niche geometries and practical data regimes under a consistent task definition.
This repository contains the benchmarking framework and code for our paper. We evaluate 16 representative algorithms spanning probabilistic models, graph neural networks, deep generative models, and foundation models. The benchmark quantifies performance across complementary axes including agreement with reference niches (accuracy), spatial structure and boundary fidelity (connectivity), biological consistency of inferred niches (composition similarity), quality of learned embeddings (silhouette score), and computational efficiency (runtime and memory).
If you find this repository or the accompanying benchmark useful for your work, please consider citing our work:
Wang, Y., Chen, Y., Yang, L., Wang, C., Cai, J., and Xin, H. Benchmarking niche identification via domain segmentation for spatial transcriptomics data. bioRxiv (2026). DOI: 10.64898/2026.02.27.708202.
We benchmarked 16 representative algorithms categorized into four methodological families.
- BayesSpace: Code | Paper: Spatial transcriptomics at subspot resolution with BayesSpace (Nature Biotechnology)
- BANKSY: Code | Paper: BANKSY unifies cell typing and tissue domain segmentation for scalable spatial omics data analysis (Nature Genetics)
- MENDER: Code | Paper: MENDER: fast and scalable tissue structure identification in spatial omics data (Nature Communications)
- SpaGCN: Code | Paper: SpaGCN: Integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network (Nature Methods)
- GraphST: Code | Paper: Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with GraphST (Nature Communications)
- STAGATE: Code | Paper: Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder (Nature Communications)
- CytoCommunity: Code | Paper: Unsupervised and supervised discovery of tissue cellular neighborhoods from cell phenotypes (Nature Methods)
- SpaceFlow: Code | Paper: Identifying multicellular spatiotemporal organization of cells with SpaceFlow (Nature Communications)
- NicheCompass: Code | Paper: Quantitative characterization of cell niches in spatially resolved omics data (Nature Genetics)
- scNiche: Code | Paper: Identification and characterization of cell niches in tissue from spatial omics data at single-cell resolution (Nature Communications)
- SEDR: Code | Paper: Unsupervised spatially embedded deep representation of spatial transcriptomics (Genome Medicine)
- STACI: Code | Paper: Graph-based autoencoder integrates spatial transcriptomics with chromatin images and identifies joint biomarkers for Alzheimer’s disease (Nature Communications)
- DeepLinc: Code | Paper: De novo reconstruction of cell interaction landscapes from single-cell spatial transcriptome data with DeepLinc (Genome Biology)
- CellCharter: Code | Paper: CellCharter reveals spatial cell niches associated with tissue remodeling and cell plasticity (Nature Genetics)
- Novae: Code | Paper: Novae: a graph-based foundation model for spatial transcriptomics data (Nature Methods)
- Nicheformer: Code | Paper: Nicheformer: a foundation model for single-cell and spatial omics (Nature Methods)
Results are quantified using the following categorical metrics:
- Ground Truth Accuracy: Measures agreement with reference annotations.
- ARI (Adjusted Rand Index)
- AMI (Adjusted Mutual Information)
- Homogeneity
- Completeness
- Macro-F1 Score
- Biological Consistency: Assesses the biological relevance of inferred niches.
- Cell Type Cosine Similarity
- Spatial Structure: Evaluates the spatial coherence of the partition.
- Spatial Connectivity
- Embedding Quality: Measures the separation and compactness of the latent representation.
- Silhouette Score
- Computational Efficiency: Assesses practical scalability.
- Runtime
- Peak Memory
The manually curated human lymph node niche annotations are available in two forms:
- A ready-to-use processed AnnData file,
lymph_node_niche_annotated.h5ad, which contains the processed lymph node data with niche annotations. - Cell-level annotation tables in
annotation/(lymph_node_annotations.tsvandlymph_node_annotations.csv) for users who want to align the annotations to the original data or to their own reprocessed AnnData object.
See annotation/README.md for the annotation schema and an example of merging the TSV/CSV annotations into an .h5ad file.
annotation/: Contains scripts and data for constructing the manually annotated high-resolution human lymph node reference and defining ground truths.benchmark/: Contains the implementation and running scripts (Jupyter Notebooks) for the 16 benchmarked methods.simulation/: Contains code for generating synthetic spatial transcriptomics data using SRTsim with controlled niche parameters.
This project is released under the MIT License. See LICENSE for details.