CHIMIYA-1: An Autoselection Foundation Model for ADMET Property Prediction, Rigorously Benchmarked Against the Therapeutics Data Commons ADMET Group
Varghese, R.; Tiwary, P.; Oswal, K.
Show abstract
Accurate, generalizable prediction of absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties remains one of the highest-leverage unsolved problems in computational drug discovery, and late-stage attrition driven by ADMET liabilities continues to be a dominant cost driver in pharmaceutical research and development. The Therapeutics Data Commons (TDC) ADMET Group has emerged as the fields most widely adopted public benchmark, comprising 22 endpoints under standardized scaffold-split evaluation. In this work we report a comprehensive evaluation of CHIMIYA-1, a proprietary autoselection foundation model developed by Covenant Biosciences, against the full TDC ADMET Group. Departing from common practice in the field, every reported score is the mean and standard deviation of five independently seeded end-to-end evaluation runs (TDCs own minimum submission standard, which we find is not met by all public leaderboard entries), and all 22 endpoints were additionally subjected to an explicit train/test structural-overlap audit prior to reporting, finding zero overlaps on any endpoint. Despite this deliberately conservative evaluation standard, CHIMIYA-1 ranks first among all publicly listed methods on four endpoints, places within the top decile of the field on twenty of twenty-two endpoints (91%), and attains a mean percentile standing near the 74th percentile across the full benchmark, with particular strength on toxicity and physicochemical-property endpoints. We further show that several top-ranked public comparators on this benchmark have been independently found to exhibit confirmed data leakage, a finding that, if anything, understates CHIMIYA-1s relative standing. All results were obtained on commodity single-GPU workstation hardware without recourse to distributed or cloud-scale training infrastructure. We discuss these results in the context of benchmark reporting norms in molecular machine learning and outline ongoing extensions, including continuous prospective-data retraining and CUDA-level throughput optimization of the underlying selection pipeline.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- EVOSYNTH: Enabling Multi-Target Drug Discovery through Latent Evolutionary Optimization and Synthesis-Aware Prioritization 95%
- Using macromolecular electron densities to improve the enrichment of active compounds in virtual screening 90%
- A network medicine framework for multi-modal data integration in therapeutic target discovery 90%
Similar papers in this journal
Similar papers in this journal
- DrugDiff - small molecule diffusion model with flexible guidance towards molecular properties 96%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 94%
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.