The differential impacts of dataset imbalance in single-cell data integration
Maan, H.; Zhang, L.; Yu, C.; Geuenich, M.; Campbell, K. R.; Wang, B.
Show abstract
Single-cell transcriptomic data measured across distinct samples has led to a surge in computational methods for data integration. Few studies have explicitly examined the common case of cell-type imbalance between datasets to be integrated, and none have characterized its impact on downstream analyses. To address this gap, we developed the Iniquitate pipeline for assessing the stability of single-cell RNA sequencing (scRNA-seq) integration results after perturbing the degree of imbalance between datasets. Through benchmarking 5 state-of-the-art scRNA-seq integration techniques in 1600 perturbed integration scenarios for a multi-sample peripheral blood mononuclear cell (PBMC) dataset, our results indicate that sample imbalance has significant impacts on downstream analyses and the biological interpretation of integration results. We observed significant variation in clustering, cell-type classification, marker gene-based annotation, and query-to-reference mapping in imbalanced settings. Two key factors were found to lead to quantitation differences after scRNA-seq integration - the cell-type imbalance within and between samples (relative cell-type support) and the relatedness of cell-types across samples (minimum cell-type center distance). To account for evaluation gaps in imbalanced contexts, we developed novel clustering metrics robust to sample imbalance, including the balanced Adjusted Rand Index (bARI) and balanced Adjusted Mutual Information (bAMI). Our analysis quantifies biologically-relevant effects of dataset imbalance in integration scenarios and introduces guidelines and novel metrics for integration of disparate datasets. The Iniquitate pipeline and balanced clustering metrics are available at https://github.com/hsmaan/Iniquitate and https://github.com/hsmaan/balanced-clustering, respectively.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 97%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 97%
- Beyond benchmarking: towards predictive models of dataset-specific single-cell RNA-seq pipeline performance 97%
Similar papers in this journal
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 97%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 96%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- GROTIA: An Interpretable Graph-Regularized Optimal TransportFramework for Diagonal Single-Cell IntegrativeAnalysis 96%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 95%
- DrivAER: Identification of driving transcriptional programs in single-cell RNA sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.