Single-cell foundation models reveal context-sensitive cancer programmes under subtype shift
Wallace, J.; Youssef, G.; Han, N.
Show abstract
Single-cell foundation models (scFMs) have shown promise as transferable representations of cellular state, but recent zero-shot evaluations suggest that they do not consistently outperform simpler baselines. We asked whether this apparent limitation reflects an intrinsic weakness of scFMs or instead the difficulty of using them without task-specific adaptation. To test this, we fine-tuned two widely used scFMs, Geneformer and scGPT, on common tumour subtypes from renal, lung, and breast cancer, and compared them with a LightGBM baseline on within-domain validation cohorts and on out-of-domain rarer, unseen cancer subtypes. Across all three organs, the models achieved near-perfect within-domain discrimination (AUROC 0.98-1.00), but differences emerged under subtype shift. On chromophobe RCC, scGPT and Geneformer achieved AUROC 0.88 and 0.92 respectively versus 0.64 for LightGBM; on SCLC, Geneformer reached 1.00 versus 0.82 for LightGBM; and on TNBC, scGPT achieved 0.80 versus 0.49 for LightGBM. To determine whether this generalisation reflected meaningful adaptation rather than arbitrary feature drift, we applied Integrated Gradients, an interpretability technique, to the fine-tuned scFMs and SHAP to LightGBM. LightGBM showed highly stable gene-importance rankings across datasets, whereas the foundation models were substantially more context-sensitive. However, this flexibility was not random: all models converged on a shared within-domain core, while scFMs acquired larger rare-subtype-specific gene sets and pathway programmes during transfer. Pathway enrichment further supported the biological relevance of these attributed genes. Together, these results suggest that fine-tuned scFMs can bridge clinically relevant domain shifts in cancer single-cell analysis and that interpretability provides a practical route to distinguishing biologically grounded adaptation from rigid reuse of training-era rules.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 96%
- Evolutionary signatures of human cancers revealed via genomic analysis of over 35,000 patients 96%
- Integrative ensemble modelling of cetuximab sensitivity in colorectal cancer PDXs 95%
Similar papers in this journal
- Impact of between-tissue differences on pan-cancer predictions of drug sensitivity 94%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 94%
- Exploring tumor-normal cross-talk with TranNet: role of the environment in tumor progression 93%
Similar papers in this journal
Similar papers in this journal
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 95%
- Benchmarking computational methods for multi-omics biomarker discovery in cancer 94%
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.