Back

Completion of the DrugMatrix Toxicogenomics Database using ToxCompl

Cong, G.; Patton, R. M.; Chao, F.; Svoboda, D. L.; Casey, W. M.; Schmitt, C. P.; Murphy, C.; Erickson, J. N.; Combs, P. A.; Auerbach, S. S.

2024-03-29 genomics
10.1101/2024.03.26.586669 bioRxiv
Show abstract

The DrugMatrix Database contains systematically generated toxicogenomics data from short-term in vivo studies for over 600 chemicals. However, most of the potential endpoints in the database are missing due to a lack of experimental measurements. We present our study on leveraging matrix factorization and machine learning methods to predict the missing values in the DrugMatrix, which includes gene expression across eight tissues on two expression platforms along with paired clinical chemistry, hematology, and histopathology measurements. One major challenge we encounter is the skewed distribution of the available measured data, in terms of both tissue sources and values. We propose a method, ToxiCompl, that applies systematic hybrid sampling guided by Bayesian optimization in conjunction with low-rank matrix factorization to recover the missing values. ToxiCompl achieves good training and validation performance from a machine learning perspective. We further conduct an in-depth validation of the predicted data from biological and toxicological perspectives with a series of analyses. These include examining the connectivity pattern of predicted gene expression responses, characterizing molecular pathway-level responses from sets of differentially expressed genes, evaluating known transcriptional biomarkers of tissue toxicity, and characterizing pre-dicted apical endpoints. Our analysis shows that the predicted differential gene expression, broadly speaking, aligns with what would be anticipated. For example, in most instances, our predicted differentially expressed gene lists offer a connectivity level comparable to that of measured data in connectivity analysis. Using Havcr1, a known transcriptional biomarker of kidney injury, we identify treatments that, based on the predicted expression data, manifest kidney toxicity in a manner that is mechanistically plausible and supported by the literature. Characterization of the predicted clinical chemistry data suggests that strong effects are relatively reliably predicted, while more subtle effects pose a greater challenge. In the case of histopathological prediction, we find a significant overprediction due to positivity bias in the measured data. Developing methods to deal with this bias is one of the areas we plan to target for future improvement. The main advantage of the ToxiCompl approach is that, in the absence of additional experimental data, it drastically extends the toxicogenomic landscape into a number of data-poor tissues, thereby allowing researchers to formulate mechanistic hypotheses about effects in tissues that have been underrepresented in the literature. All measured and predicted DrugMatrix data (i.e., gene expression, clinical chemistry, hematology, and histopathology) are available to the public through an intuitive GUI interface that allows for data retrieval, gene set analysis and high dimensional visualization of gene expression similarity (https://rstudio.niehs.nih.gov/complete_drugmatrix/).

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 2%
15.3%
2
PLOS ONE
5266 papers in training set
Top 17%
10.7%
3
Addiction Neuroscience
17 papers in training set
Top 0.1%
6.3%
4
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
4.9%
5
Scientific Reports
3612 papers in training set
Top 20%
4.9%
6
Bioinformatics
1204 papers in training set
Top 5%
3.3%
7
BMC Bioinformatics
457 papers in training set
Top 3%
3.3%
8
Journal of Biomedical Informatics
47 papers in training set
Top 0.5%
3.3%
50% of probability mass above
9
BMC Genomics
406 papers in training set
Top 2%
3.1%
10
Mathematical Biosciences
49 papers in training set
Top 0.4%
2.8%
11
Archives of Clinical and Biomedical Research
28 papers in training set
Top 0.2%
2.5%
12
Bioinformatics Advances
203 papers in training set
Top 2%
2.1%
13
International Journal of Molecular Sciences
494 papers in training set
Top 7%
1.9%
14
The Pharmacogenomics Journal
11 papers in training set
Top 0.1%
1.9%
15
JAMIA Open
42 papers in training set
Top 0.9%
1.7%
16
GigaScience
212 papers in training set
Top 2%
1.7%
17
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.7%
18
Artificial Intelligence in Medicine
17 papers in training set
Top 0.4%
1.3%
19
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
20
PLOS Global Public Health
344 papers in training set
Top 7%
1.1%
21
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.1%
22
Frontiers in Genetics
230 papers in training set
Top 4%
1.0%
23
G3: Genes, Genomes, Genetics
252 papers in training set
Top 4%
1.0%
24
eLife
5828 papers in training set
Top 61%
1.0%
25
Clinical Pharmacology & Therapeutics
25 papers in training set
Top 0.3%
1.0%
26
Cancers
213 papers in training set
Top 5%
0.9%
27
Heliyon
152 papers in training set
Top 7%
0.9%
28
Toxicological Sciences
41 papers in training set
Top 0.5%
0.9%
29
Frontiers in Psychiatry
87 papers in training set
Top 2%
0.9%
30
BioData Mining
22 papers in training set
Top 1.0%
0.6%