Robust annotation and discovery of novel cell types in single-cell ATAC-seq data through cross-modal reference alignment.
CARA achieves accurate cross-modal cell type annotation and novel cell type discovery in single-cell ATAC-seq data, outperforming baseline methods across diverse datasets.
- Why it matters: Precise cell type identification in scATAC-seq is hindered by data sparsity, high dimensionality, limited labeled references, and batch effects, impeding biological insights.
- What they did: A cross-omics Bayesian framework was developed that leverages scRNA-seq data through pretraining and semi-supervised learning, incorporating distribution alignment and dynamic class weighting to annotate and detect new cell types.
- The result: CARA consistently enhances annotation accuracy, preserves lineage structures, and identifies rare or novel populations, enabling deeper understanding of cell regulation and facilitating extension to other modalities like DNA methylation.