Reliable single-cell perturbations explain and improve model performance
Wang, X.; Kuipers, J.; Hugi, F.; Platt, R. J.; Beerenwinkel, N.
Show abstract
Predicting single-cell transcriptional responses to perturbations is central to building the virtual cell, yet recent benchmarks show that simple baseline methods often outperform complex models, and model comparisons depend on the evaluation metric. Most studies assume that preprocessed RNA sequencing data are reliable ground truth for both training and evaluation. Here, we test this assumption by measuring the reliability of perturbations and their alignment with shared perturbation responses, classifying each perturbation as specific, shared, or unreliable. Among 7,170 perturbations from 29 datasets, 65% are unreliable, 11% shared, and 24% specific. Applying these quality labels to published benchmarks shows that model comparisons depend on perturbation quality. Training with reliable perturbations alone matches or outperforms full-data performance while using 55% of all training perturbations. Our framework also enables prospective experimental design: for most perturbations, a 28-cell pilot experiment accurately predicts how many cells a full screen needs to be reliable.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
- Towards inferring causal gene regulatory networks from single cell expression measurements 94%
- Comprehensive prediction of robust synthetic lethality between paralog pairs in cancer cell lines 94%
Similar papers in this journal
- EvoRMD: Integrating Biological Context and Evolutionary RNA Language Models for Interpretable Prediction of RNA Modifications 94%
- geneBasis: an iterative approach for unsupervised selection of targeted gene panels from scRNA-seq. 94%
- scAlign: a tool for alignment, integration and rare cell identification from scRNA-seq data 94%
Similar papers in this journal
- EternaBrain: Automated RNA design through move sets from an Internet-scale RNA videogame 95%
- Discovering functional sequences with RELICS, an analysis method for tiling CRISPR screens 94%
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.