Back

Assessment Of Alphafold Protein Models For Small-Molecule Ligand Docking

Maidanik, H.; Lazou, M.; Bajaj, R.; Sarwar, R.; Vajda, S.; Joseph-McCarthy, D.

2026-01-04 bioinformatics
10.64898/2026.01.04.697577 bioRxiv
Show abstract

Molecular docking is a powerful computational tool for predicting protein-ligand interactions, widely employed in drug discovery. However, its effectiveness is often constrained by the availability of experimentally resolved X-ray protein structures, a process that is both time consuming and resource-intensive. AlphaFold (AF), a deep learning method, offers an efficient alternative by predicting high-accuracy 3D protein structures directly from amino acid sequences. This study assesses the utility of AF-generated protein models for fragment and larger ligand docking with Glide, a widely used docking approach. The docking workflow is evaluated in an unbiased manner by carrying out binding site identification with FTMap, a binding hot spot prediction software. We show that fragment docking to AF models outperforms docking to the respective unbound protein crystal structures, and performs comparably to docking to the corresponding ligand bound structures when using an unbiased approach. Leveraging computational efficiency of AF model generation, we also employ ensembles of AF models to incorporate protein flexibility. Results show that docking to AF ensembles improves larger-ligand docking compared to docking to singular AF models and outperforms docking to unbound structures. The results provide insights into the effectiveness of integrating AF protein models into docking procedures, highlighting the potential for streamlining computational drug discovery processes. STATEMENT OF SIGNIFICANCEThis work addresses a critical bottleneck in computational drug discovery by demonstrating that AlphaFold (AF) models can serve as an alternative or complement to experimental structures for molecular docking. Specifically, a systematic study assessing Glide docking to rigid AF protein models and ensembles of models compared to experimentally determined ligand-bound and unbound protein X-ray structures was performed. The evaluation employs an unbiased methodology using FTMap-identified binding sites, eliminating the need for prior knowledge of the native ligand binding location. Additionally, protein flexibility is incorporated through a multiseed ensemble approach that generates a conformational ensembles of AF models at minimal computational cost, improving the docking accuracy without the need for ligand-bound templates.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.