Efficient and scalable integration of single-cell data using domain-adversarial and variational approximation
Hu, J.; Zhong, Y.; Shang, X.
Show abstract
Single-cell data provides us new ways of discovering biological truth at the level of individual cells, such as identification of cellular sub-populations and cell development. With the development of single-cell sequencing technologies, a key analytical challenge is to integrate these data sets to uncover biological insights. Here, we developed a domain-adversarial and variational approximation framework, DAVAE, to integrate multiple single-cell data across samples, technologies and modalities without any post hoc data processing. We fit normalized gene expression into a non-linear model, which transforms a latent variable of a lower-dimension into expression space with a non-linear function, a KL regularizier and a domain-adversarial regularizer. Results on five real data integration applications demonstrated the effectiveness and scalability of DAVAE in batch-effect removing, transfer learning, and cell type predictions for multiple single-cell data sets across samples, technologies and modalities. DAVAE was implemented in the toolkit package "scbean" in the pypi repository, and the source code can be also freely accessible at https://github.com/jhu99/scbean.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 98%
- A Robust and Scalable Graph Neural Network for Accurate Single Cell Classification 98%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 98%
Similar papers in this journal
- scGCN: a Graph Convolutional Networks Algorithm for Knowledge Transfer in Single Cell Omics 98%
- Deep autoencoder for interpretable tissue-adaptive deconvolution and cell-type-specific gene analysis 98%
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 98%
Similar papers in this journal
- BERMUDA: A novel deep transfer learning method for single-cell RNA sequencing batch correction reveals hidden high-resolution cellular subtypes 98%
- scINSIGHT for interpreting single-cell gene expression from biologically heterogeneous data 97%
- Learning latent embedding of multi-modal single cell data and cross-modality relationship simultaneously 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.