Uncertainty-aware synthetic lethality prediction with pretrained foundation models
Hua, K.; Haber, E.; Ma, J.
Show abstract
Synthetic lethality (SL) offers a promising paradigm for targeted cancer therapy, yet experimental identification of SL gene pairs remains costly, context-dependent, and biased toward well-studied genes. Existing computational approaches often rely on curated protein-protein interaction (PPI) networks and Gene Ontology (GO) annotations, which limit their ability to generalize to novel genes. Here we introduce CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW, a two-stage, graph-free framework that leverages pretrained biological foundation models to predict SL pairs with calibrated uncertainty. In Stage 1, we apply a pretrained single-cell foundation model to bulk RNA-seq profiles of cancer cell lines to obtain context-aware embeddings and perform in silico gene knockouts to generate delta embeddings. These perturbation signals are further conditioned on a data-driven gene prior and supervised with CRISPR viability readouts to learn knockout-aware viability embeddings. In Stage 2, we derive pairwise features from these embeddings and train a lightweight classifier to distinguish SL from non-SL pairs. To enable reliable experimental prioritization, CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW incorporates conformal prediction, producing calibrated and interpretable prediction sets that highlight high-confidence SL candidates. Across two evaluation settings, including zero-shot generalization to unseen gene pairs and to unseen genes, ablation analyses show that viability pretraining and the gene prior substantially improve performance while avoiding reliance on PPI and GO features. CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW therefore transforms pretrained biological representations into practical, uncertainty-aware hypotheses that support robust and scalable discovery of therapeutic targets.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 97%
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 97%
- scDisInFact: disentangled learning for integration and prediction of multi-batch multi-condition single-cell RNA-sequencing data 96%
Similar papers in this journal
Similar papers in this journal
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 96%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 95%
- An adversarial scheme for integrating multi-modal data on protein function 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.