Back

Leveraging Large Language Models for Literature-Driven Prioritization of Protein Binding Pockets

Stratiichuk, R.; Melnychenko, M.; Koleiev, I.; Voitsitskyi, T.; Vladyslav, H.; Shevchuk, N.; Osrovsky, Z.; Bdzhola, V.; Yesylevskyy, S.; Starosyla, S.; Nafiiev, A.

2025-05-15 biophysics
10.1101/2025.05.13.653394 bioRxiv
Show abstract

We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with Large Language Models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLMs performance in paper relevance assessment and pocket extraction. The extracted information is then mapped onto protein structures and used to filter and merge the geometry-based predictions, generating a refined volumetric representation of biologically relevant pockets. This hybrid pipeline offers an efficient, accurate and automated method for identifying functional binding pockets, addressing a significant bottleneck in the high-throughput drug discovery workflows. The developed benchmark dataset and methodology are freely available at https://github.com/MelnychenkoM/LLM-benchmark-dataset.

Published in Bioinformatics (predicted rank #3) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.