SmartImpute: A Targeted Imputation Framework for Single-cell Transcriptome Data
Yao, S.; Yu, X.; Wang, X.
Show abstract
Single-cell RNA sequencing (scRNA-seq) has revolutionized our understanding of cellular heterogeneity and tissue transcriptomic complexity. However, the high frequency of dropout events in scRNA-seq data complicates downstream analyses such as cell type identification and trajectory inference. Existing imputation methods address the dropout problem but face limitations such as high computational cost and risk of over-imputation. We present SmartImpute, a novel computational framework designed for targeted imputation of scRNA-seq data. SmartImpute focuses on a predefined set of marker genes, enhancing the biological relevance and computational efficiency of the imputation process while minimizing the risk of model misspecification. Utilizing a modified Generative Adversarial Imputation Network architecture, SmartImpute accurately imputes the missing gene expression and distinguishes between true biological zeros and missing values, preventing overfitting and preserving biologically relevant zeros. To ensure reproducibility, we also provide a function based on the GPT4 model to create target gene panels depending on the tissue types and research context. Our results, based on scRNA-seq data from head and neck squamous cell carcinoma and human bone marrow, demonstrate that SmartImpute significantly enhances cell type annotation and clustering accuracy while reducing computational burden. Benchmarking against other imputation methods highlights SmartImputes superior performance in terms of both accuracy and efficiency. Overall, SmartImpute provides a lightweight, efficient, and biologically relevant solution for addressing dropout events in scRNA-seq data, facilitating deeper insights into cellular heterogeneity and disease progression. Furthermore, SmartImputes targeted approach can be extended to spatial omics data, which also contain many missing values.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scPDA: Denoising Protein Expression in Droplet-Based Single-Cell Data 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- BERMUDA: A novel deep transfer learning method for single-cell RNA sequencing batch correction reveals hidden high-resolution cellular subtypes 96%
Similar papers in this journal
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 96%
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 95%
- Single-cell Multi-omics Integration for Unpaired Data by a Siamese Network with Graph-based Contrastive Loss 95%
Similar papers in this journal
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 97%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 96%
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.