Fast clustering and cell-type annotation of scATAC data using pre-trained embeddings
LeRoy, N. J.; Smith, J. P.; Zheng, G.; Rymuza, J.; Gharavi, E.; Brown, D. E.; Zhang, A.; Sheffield, N. C.
Show abstract
MotivationData from the single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) is now widely available. One major computational challenge is dealing with high dimensionality and inherent sparsity, which is typically addressed by producing lower-dimensional representations of single cells for downstream clustering tasks. Current approaches produce such individual cell embeddings directly through a one-step learning process. Here, we propose an alternative approach by building embedding models pre-trained on reference data. We argue that this provides a more flexible analysis workflow that also has computational performance advantages through transfer learning. ResultsWe implemented our approach in scEmbed, an unsupervised machine learning framework that learns low-dimensional embeddings of genomic regulatory regions to represent and analyze scATAC-seq data. scEmbed performs well in terms of clustering ability and has the key advantage of learning patterns of region co-occurrence that can be transferred to other, unseen datasets. Moreover, pre-trained models on reference data can be exploited to build fast and accurate cell-type annotation systems without the need for other data modalities. scEmbed is implemented in Python and it is available to download from GitHub. We also make our pre-trained models available on huggingface for public use. AvailabilityscEmbed is open source and available at https://github.com/databio/geniml. Pre-trained models from this work can be obtained on huggingface: https://huggingface.co/databio.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CMOT: Cross Modality Optimal Transport for multimodal inference 97%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 97%
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 97%
Similar papers in this journal
Similar papers in this journal
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 97%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 96%
- Learning interpretable representations of single-cell multi-omics data with multi-output Gaussian Processes 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.