Back

Cross-chemical and cross-species toxicity prediction: benchmarkingand a novel 3D-structure-based deep learning model

Yuan, R.; Shaw, J.; Tang, H.; Ye, Y.

2025-11-26 bioinformatics
10.1101/2025.11.24.690199 bioRxiv
Show abstract

Prediction of a compounds toxicity is a key step toward realizing animal-free testing of chemical compounds. Recent advances have yielded significant progress in computational toxicity prediction, including machine learning methods that utilize chemical fingerprints and deep learning-based latent representations. However, challenges remain, primarily due to the lack of clean training datasets and the inconsistent model performance. To address these challenges, we curated a comprehensive dataset of aquatic toxicity from seven data sources, which contains 50,603 records for 5,889 compounds across 2,285 different species, much larger than similar datasets used in previous studies. We also developed tox-learn, a Python library featuring tools for automated dataset cleaning, machine learning methods and performance evaluation. The library places special emphasis on avoiding overestimation of prediction accuracy caused by improper train-test data splitting. Based on this toolbox, we benchmarked various predictive models using different train-test splitting strategies on the curated dataset. Our results showed that the choice of machine learning method, molecular fingerprint, and train-test splitting strategy all significantly affect performance. We demonstrated that incorporating species information generally improved predictions, although the degree of improvement depended on how this information was represented. In addition, we developed a new 3D structure-based deep-learning model, 3DMol-Tox, which achieves regression accuracy comparable to the best 2D-structure based model (GPBoost) while exhibiting consistently higher within-one-bin (W1B) classification accuracy. Finally, we analyzed the impact of different train-test splitting strategies and provide recommendations based on our benchmarking, such as using structure-aware splitting to mitigate information leakage, a common issue that inflates reported model performance.

Published in Environmental Toxicology and Chemistry · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of Cheminformatics
29 papers in training set
Top 0.1%
18.7%
2
Bioinformatics
1204 papers in training set
Top 2%
9.9%
3
PLOS Computational Biology
1863 papers in training set
Top 4%
8.0%
4
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.6%
8.0%
5
PLOS ONE
5266 papers in training set
Top 27%
5.6%
50% of probability mass above
6
BMC Bioinformatics
457 papers in training set
Top 2%
4.4%
7
Scientific Reports
3612 papers in training set
Top 29%
3.6%
8
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1%
3.3%
9
Communications Chemistry
48 papers in training set
Top 0.3%
2.7%
10
Nature Communications
5641 papers in training set
Top 39%
2.5%
11
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.5%
12
Chemical Research in Toxicology
10 papers in training set
Top 0.1%
2.1%
13
Bioinformatics Advances
203 papers in training set
Top 2%
2.1%
14
Scientific Data
209 papers in training set
Top 1%
1.7%
15
GigaScience
212 papers in training set
Top 3%
1.3%
16
Metabolites
53 papers in training set
Top 0.6%
1.3%
17
Molecules
39 papers in training set
Top 1.0%
1.1%
18
Cell Systems
201 papers in training set
Top 4%
1.0%
19
Science Advances
1243 papers in training set
Top 28%
1.0%
20
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.2%
0.9%
21
Advanced Science
286 papers in training set
Top 9%
0.9%
22
npj Digital Medicine
118 papers in training set
Top 3%
0.9%
23
Nature Protocols
33 papers in training set
Top 0.5%
0.9%
24
Patterns
78 papers in training set
Top 3%
0.9%
25
Archives of Toxicology
18 papers in training set
Top 0.3%
0.9%
26
ACS Omega
105 papers in training set
Top 4%
0.6%