Gene-embedding-based prediction and functional evaluation of perturbation expression responses with PRESAGE
Littman, R.; Levine, J.; Maleki, S.; Lee, Y.; Ermakov, V.; Qiu, L.; Wu, A.; Huang, K.; Lopez, R.; Scalia, G.; Biancalani, T.; Richmond, D.; Regev, A.; Hütter, J.-C.
Show abstract
Understanding the impact of genetic perturbations on cellular behavior is crucial for biological research, but comprehensive experimental mapping remains infeasible. We introduce PRESAGE (Perturbation Response EStimation with Aggregated Gene Embeddings), a simple, modular, and interpretable framework that predicts perturbation-induced expression changes by integrating diverse knowledge sources via gene embeddings. PRESAGE transforms gene embeddings through an attention-based model to predict perturbation expression outcomes. To assess model performance, we introduce a comprehensive evaluation suite with novel functional metrics that move beyond traditional regression tasks, including measures of accuracy in effect size prediction, in identifying perturbations with similar expression profiles (phenocopy), and in prediction of perturbations with the strongest impact on specific gene set scores. PRESAGE outperforms existing methods in both classical regression metrics and our novel functional evaluations. Through ablation studies, we demonstrate that knowledge source selection is more critical for predictive performance than architectural complexity, with cross-system Perturb-seq data providing particularly strong predictive power. We also find that performance saturates quickly with training set size, suggesting that experimental design strategies might benefit from collecting sparse perturbation data across multiple biological systems rather than exhaustive profiling of individual systems. Overall, PRESAGE establishes a robust framework for advancing perturbation response prediction and facilitating the design of targeted biological experiments, significantly improving our ability to predict cellular responses across diverse biological systems.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Iterative deep learning-design of human enhancers exploits condensed sequence grammar to achieve cell type-specificity 96%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 95%
- Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.