Quantifying Cell-type-specific Differences of Single-cell Datasets using UMAP and SHAP
Lim, H. S.; Qiu, P.
Show abstract
With the rapid advances in single-cell profiling technologies, larger-scale investigations that require comparisons of multiple single-cell datasets can lead to novel findings. Specifically, quantifying cell-type-specific responses to different conditions across single-cell datasets could be useful in understanding how the difference in conditions is induced at a cellular level. Here we present a computational pipeline that quantifies the cell-type-specific differences and identifies genes responsible for the differences. We quantify differences observed in a low-dimensional UMAP space as a proxy for the difference present in the high-dimensional space and use SHAP to quantify genes driving the differences. Here we applied our algorithm to the Iris flower dataset, scRNA-seq dataset, and mass cytometry dataset, and demonstrate that it can robustly quantify the cell-type-specific differences and it can also identify genes that are responsible for the differences.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 97%
- DGCyTOF: deep learning with graphic cluster visualization to predict cell types of single cell mass cytometry data 96%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 96%
Similar papers in this journal
Similar papers in this journal
- Coffee: Consensus Single Cell-Type Specific Inference For Gene Regulatory Networks 96%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 96%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.