Zero-shot biological reasoning with open-weights large language models reproduces CRISPR screen based prediction of synthetic lethal interactions.
Prosz, A. G.; Sztupinszki, Z.; Diossy, M.; Zimon, B.; Csabai, I. G.; Szallasi, Z.
Show abstract
Identifying clinically relevant synthetic lethal interactions has great potential for uncovering novel therapeutic vulnerabilities in cancer. Current approaches rely on machine learning models that estimate probabilities of synthetic lethal interactions, without supplying explicit knowledge of the underlying biology and lack the human-readable interpretation leading to the prediction. Large Language Models (LLMs) represent a new class of tools capable of reasoning and leveraging extensive biological knowledge acquired from relevant literature during their pretraining. Here, we tested multiple open-weight LLMs for their ability to predict known and novel synthetic lethal interactions. We found that most of the tested models were better at reconstructing the results of three known genome-wide CRISPR knockout screens than random chance, while observed that their performance was related to the parameter-size of the model, and on average benefited little from additional pathway and genetic information apart from what they already possess when estimating the likelihood of a synthetic lethal relationship. After selecting the best-performing and most computationally efficient model for our use case (Qwen2.5-32B-Instruct, 0.715 AUROC), we performed an in silico screen of 398,277 gene pairs from 893 clinically relevant genes. Our goal was to highlight the potential of open-weights LLMs as scalable, context-aware prioritization tools for synthetic lethal interactions, and to lay the groundwork for predicting higher-order genetic interactions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 94%
- Application of Modular Response Analysis to Medium- to Large-Size Biological Systems 94%
Similar papers in this journal
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 95%
- A benchmark of computational methods for correcting biases of established and unknown origin in CRISPR-Cas9 screening data 95%
- Network Propagation-based Prioritization of Long Tail Genes in 17 Cancer Types 94%
Similar papers in this journal
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- Beyond synthetic lethality in large-scale metabolic and regulatory network models via genetic minimal intervention sets 94%
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 93%
Similar papers in this journal
- A multi-task domain-adapted model to predict chemotherapy response from mutations in recurrently altered cancer genes 93%
- A Highly-Efficient, Scalable Pipeline for Fixed Feature Extraction from Large-Scale High-Content Imaging Screens 92%
- CELLoGeNe - an Energy Landscape Framework for Logical Networks Controlling Cell Decisions 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.