Comparison of marker selection methods for high throughput scRNA-seq data
Vargo, A. H.; Gilbert, A. C.
Show abstract
Here, we evaluate the performance of a variety of marker selection methods on scRNA-seq UMI counts data. We test on an assortment of experimental and synthetic data sets that range in size from several thousand to one million cells. In addition, we propose several performance measures for evaluating the quality of a set of markers when there is no known ground truth. According to these metrics, most existing marker selection methods show similar performance on experimental scRNA-seq data; thus, the speed of the algorithm is the most important consid-eration for large data sets. With this in mind, we introduce RO_SCPCAPANKC_SCPCAPCO_SCPCAPORRC_SCPCAP, a fast marker selection method with strong mathematical underpinnings that takes a step towards sensible multi-class marker selection.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 96%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
- Assessing the Performance of Methods for Cell Clustering from Single-cell DNA Sequencing Data 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MOSCATO: A Supervised Approach for Analyzing Multi-Omic Single-Cell Data 95%
- Genomic prediction using machine learning: A comparison of the performance of regularized regression, ensemble, instance-based and deep learning methods on synthetic and empirical data 94%
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.