Exposing the Molecular Reaction Blind Spots of LLMs with PathwayQA
Nayar, G.; Carpenter, K. A.; Smith, D. A.; Xiong, B.; Altman, R. B.
Show abstract
Proteins mediate a large portion of cellular activity, and understanding protein pathways can yield novel biological insights. Large language models (LLMs) have become increasingly adept at performing inference tasks across different fields of science and engineering. These models could facilitate the analysis of protein networks and help generate hypotheses about protein interactions in a scalable and accessible manner. However, the performance of LLMs in inferring protein-mediated biochemical reactions remains understudied. Here, we evaluate nine LLMs in reasoning over protein pathways included in the curated Reactome database. We find that all nine models struggle to infer products of a reaction when given reactants and enzymes. GPT-4o mini performed the best with a median recovery score of 0.6667, but no model surpassed the baseline strategy of parroting reactants back as predicted products. Most LLMs also performed poorly when inferring whether a protein pathway is associated with a human disease, with an average accuracy of 0.5980. DeepSeek 7B Chat performed the best with an accuracy of 0.9100. This study highlights an area where LLMs still struggle to make correct inferences and provides an opportunity for further work in developing biological LLMs. We also provide a novel question-answer dataset, PathwayQA, which is based on the Reactome database. PathwayQA can be used to benchmark and improve model performance on reasoning over protein-interaction networks. PathwayQA is available at https://github.com/Helix-Research-Lab/PathwayQA.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MENDELSEEK: An algorithm that predicts Mendelian Genes and elucidates what makes them special 94%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 93%
- Non-linear Archetypal Analysis of Single-cell RNA-seq Data by Deep Autoencoders 93%
Similar papers in this journal
- A versatile information retrieval framework for evaluating profile strength and similarity 94%
- Triple-effect correction for Cell Painting data with contrastive and domain-adversarial learning 94%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 93%
Similar papers in this journal
- Cracking the black box of deep sequence-based protein-protein interaction prediction 95%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.