Deep learning does not outperform classical machine learning for cell-type annotation
Köhler, N. D.; Büttner, M.; Andriamanga, N.; Theis, F. J.
Show abstract
Deep learning has revolutionized image analysis and natural language processing with remarkable accuracies in prediction tasks, such as image labeling and semantic segmentation or named-entity recognition and semantic role labeling. Specifically, the combination of algorithmic and hardware advances with the appearance of large and well-labeled datasets has led up to seminal contributions in these fields. The emergence of large amounts of data from single-cell RNA-seq and the recent global effort to chart all cell types in the Human Cell Atlas has attracted an interest in deep-learning applications. However, all current approaches are unsupervised, i.e., learning of latent spaces without using any cell labels, even though supervised learning approaches are often more powerful in feature learning and the most popular approach in the current AI revolution by far. Here, we ask why this is the case. In particular we ask whether supervised deep learning can be used for cell annotation, i.e. to predict cell-type labels from single-cell gene expression profiles. After evaluating 10 classification methods across 14 datasets, we notably find that deep learning does not outperform classical machine-learning methods in the task. Thus, cell-type prediction based on gene-signature derived cell-type labels is potentially too simplistic a task for complex non-linear methods, which demands better labels of functional single-cell readouts.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting cell types with supervised contrastive learning on cells and their types 95%
- Correspondence analysis for dimension reduction, batch integration, and visualization of single-cell RNA-seq data 95%
- Self-supervision advances morphological profiling by unlocking powerful image representations 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.