Have protein-ligand co-folding methods moved beyond memorisation?
Skrinjar, P.; Eberhardt, J.; Durairaj, J.; Schwede, T.
10.1101/2025.02.03.636309 bioRxivShow abstract
Deep learning has driven major breakthroughs in protein structure prediction, however the next critical advance is accurately predicting how proteins interact with small molecule ligands, to enable real-world applications such as drug discovery. Recent cofolding methods aim to address this challenge, but evaluating their performance has been inconclusive due to the lack of relevant bench-marking datasets. Here we present a comprehensive evaluation of four leading all-atom cofolding methods using our newly introduced benchmark dataset Runs N Poses, which comprises 2,600 high-resolution protein-ligand systems released after the training cutoff used by these methods. We demonstrate that current cofolding approaches largely memorise ligand poses from their training data, hindering their use for de novo drug design. With this assessment and benchmark dataset, we aim to accelerate progress in the field by allowing for a more realistic assessment of the current state-of-the-art deep learning methods for predicting protein-ligand interactions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
- Prioritizing virtual screening with interpretable interaction fingerprints 96%
- Dataset Augmentation Allows Deep Learning-Based Virtual Screening To Better Generalize To Unseen Target Classes, And Highlight Important Binding Interactions 96%
Similar papers in this journal
- To pack or not to pack: revisiting protein side-chain packing in the post-AlphaFold era 97%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 94%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.