Leveraging Large Language Models for Literature-Driven Prioritization of Protein Binding Pockets
Stratiichuk, R.; Melnychenko, M.; Koleiev, I.; Voitsitskyi, T.; Vladyslav, H.; Shevchuk, N.; Osrovsky, Z.; Bdzhola, V.; Yesylevskyy, S.; Starosyla, S.; Nafiiev, A.
Show abstract
We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with Large Language Models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLMs performance in paper relevance assessment and pocket extraction. The extracted information is then mapped onto protein structures and used to filter and merge the geometry-based predictions, generating a refined volumetric representation of biologically relevant pockets. This hybrid pipeline offers an efficient, accurate and automated method for identifying functional binding pockets, addressing a significant bottleneck in the high-throughput drug discovery workflows. The developed benchmark dataset and methodology are freely available at https://github.com/MelnychenkoM/LLM-benchmark-dataset.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast Local Alignment of Protein Pockets (FLAPP): A system-compiled program for large-scale binding site alignment 97%
- DeepFrag: An Open-Source Browser App for Deep-Learning Lead Optimization 96%
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
Similar papers in this journal
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 96%
- NRGSuite-Qt: A PyMOL plugin for high-throughput virtual screening, molecular docking, normal-mode analysis, the study of molecular interactions and the detection of binding-site similarities 96%
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.