Critical Assessment of ML models for ADMET Prediction in TDC leaderboards
Koleiev, I.; Stratiichuk, R.; Shevchuk, N.; Melnychenko, M.; Nyporko, O.; Todoryshyn, D.; Husak, V.; Starosyla, S.; Yesylevskyy, S. O.; Nafiiev, A.
Show abstract
In this work we performed a critical assessment of the benchmarking procedures used in Therapeutics Data Commons (TDC) ADMET leaderboards, focusing on reproducibility, robustness against data leakage, and signs of test-set overfitting across all 22 TDC ADMET endpoints. For each endpoint, the top 3 leaderboard models were screened with a unified protocol: execution environment reproducibility check, data leakage assessment, verification of hyperparameter optimisation practices, and final re-evaluation of results and TDC ranking. Only 3 methods (CaliciBoost, MapLight, MapLight+GNN) passed all checks and showed overall reproducible performance, whereas most of top-ranked models exhibited unavailable code, non-reproducible execution environments, runtime incompatibilities, or various methodological flaws. In particular, we identified direct or indirect data leakages in MiniMol, GradientBoost and XGBoost models. We also used our in-house models based on the Mol2Vec architecture to investigate the consequences of deliberately overfitting on the TDC test set. It is shown that deliberate or accidental tuning on the public test set may lead to significant inflation of the model metrics and leaderboard position. Our results emphasize the urgent need for better public ADMET benchmarks with the hidden test sets, strict dataset versioning and model submission with standardized inference environments.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 95%
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 94%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 93%
Similar papers in this journal
- Controlling astrocyte-mediated synaptic pruning signals for schizophrenia drug repurposing with Deep Graph Networks 95%
- Protein Domain-Based Prediction of Compound-Target Interactions and Experimental Validation on LIM Kinases 93%
- dGPredictor: Automated fragmentation method for metabolic reaction free energy prediction and de novo pathway design 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.