Domain-adaptation deep learning models do not outperform simple baseline models in single-cell anti-cancer drug sensitivity prediction
Esteban-Medina, M.; Bohl, M.; Beerenwinkel, N.; Lenhof, K.
Show abstract
Tumor drug response is profoundly shaped by cellular heterogeneity, making single-cell resolution essential for precision oncology. While drug-response labels are abundant for cell lines at bulk resolution, translating these predictive models to the single-cell level requires effective domain adaptation strategies. Motivated by advances in computer vision, recent deep-learning domain adaptation methods promise to transfer knowledge from bulk (source) to single-cell (target) data with-out the need for target labels. However, their true translational utility remains unclear due to a lack of rigorous evaluation against non-adaptive baselines across diverse biological and technical contexts. Here, we present a comprehensive benchmark comparing four representative domain adaptation methods against two simple gradient boosting baseline methods. Through systematic evaluation across 19 single-cell datasets and 10 drugs, we show that none of the complex adaptation methods outperforms the simpler baselines. By analyzing the drivers of model performance, we find that target-informed hyperparameter tuning and sparse label supervision are the principal sources of prediction gain. Our study reveals that current approaches fail to bridge the bulk-to-single-cell conceptual shift and provides a unified codebase and comprehensive data collection to facilitate robust model comparisons. By enabling transparent evaluation and robust benchmarking against simple models, this resource aims to accelerate future developments in translational pharmacogenomics.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 96%
- comboFM: leveraging multi-way interactions for systematic prediction of drug combination effects 96%
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 95%
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 95%
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 94%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 94%
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- The Specious Art of Single-Cell Genomics 95%
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.