PLINDER: The protein-ligand interactions dataset and evaluation resource
Durairaj, J.; Adeshina, Y.; Cao, Z.; Zhang, X.; Oleinikovas, V.; Duignan, T.; McClure, Z.; Robin, X.; Kovtun, D.; Rossi, E.; Zhou, G.; Veccham, S.; Isert, C.; Peng, Y.; Sundareson, P.; Akdel, M.; Corso, G.; Stärk, H.; Carpenter, Z.; Bronstein, M.; Kucukbenli, E.; Schwede, T.; Naef, L.
Show abstract
Protein-ligand interactions (PLI) are foundational to small molecule drug design. With computational methods striving towards experimental accuracy, there is a critical demand for a well-curated and diverse PLI dataset. Existing datasets are often limited in size and diversity, and commonly used evaluation sets suffer from training information leakage, hindering the realistic assessment of method generalization capabilities. To address these shortcomings, we present PLIN-DER, the largest and most annotated dataset to date, comprising 449,383 PLI systems, each with over 500 annotations, similarity metrics at protein, pocket, interaction and ligand levels, and paired unbound (apo) and predicted structures. We propose an approach to generate training and evaluation splits that minimizes task-specific leakage and maximizes test set quality, and compare the resulting performance of DiffDock when retrained with different kinds of splits.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 97%
- Prioritizing virtual screening with interpretable interaction fingerprints 95%
- Dataset Augmentation Allows Deep Learning-Based Virtual Screening To Better Generalize To Unseen Target Classes, And Highlight Important Binding Interactions 95%
Similar papers in this journal
- Robustly interrogating machine learning based scoring functions: what are they learning? 97%
- InterPepScore: A Deep Learning Score for Improving the FlexPepDock Refinement Protocol 96%
- QDeep: distance-based protein model quality estimation by residue-level ensemble error classifications using stacked deep residual neural networks 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.