Leveraging High-Throughput Screening Data and Conditional Generative Adversarial Networks to Advance Predictive Toxicology
Green, A. J.; Mohlenkamp, M. J.; Das, J.; Chaudhari, M.; Truong, L.; Tanguay, R. L.; Reif, D. M.
Show abstract
There are currently 85,000 chemicals registered with the Environmental Protection Agency (EPA) under the Toxic Substances Control Act, but only a small fraction have measured toxicological data. To address this gap, high-throughput screening (HTS) methods are vital. As part of one such HTS effort, embryonic zebrafish were used to examine a suite of morphological and mortality endpoints at six concentrations from over 1,000 unique chemicals found in the ToxCast library (phase 1 and 2). We hypothesized that by using a conditional Generative Adversarial Network (cGAN) and leveraging this large set of toxicity data, plus chemical structure information, we could efficiently predict toxic outcomes of untested chemicals. CAS numbers for each chemical were used to generate textual files containing three-dimensional structural information for each chemical. Utilizing a novel method in this space, we converted the 3D structural information into a weighted set of points while retaining all information about the structure. In vivo toxicity and chemical data were used to train two neural network generators. The first used regression (Go-ZT) while the second utilized cGAN architecture (GAN-ZT) to train a generator to produce toxicity data. Our results showed that both Go-ZT and GAN-ZT models produce similar results, but the cGAN achieved a higher sensitivity (SE) value of 85.7% vs 71.4%. Conversely, Go-ZT attained higher specificity (SP), positive predictive value (PPV), and Kappa results of 67.3%, 23.4%, and 0.21 compared to 24.5%, 14.0%, and 0.03 for the cGAN, respectively. By combining both Go-ZT and GAN-ZT, our consensus model improved the SP, PPV, and Kappa, to 75.5%, 25.0%, and 0.211, respectively, resulting in an area under the receiver operating characteristic (AUROC) of 0.663. Considering their potential use as prescreening tools, these models could provide in vivo toxicity predictions and insight into untested areas of the chemical space to prioritize compounds for HT testing. SummaryA conditional Generative Adversarial Network (cGAN) can leverage a large chemical set of experimental toxicity data plus chemical structure information to predict the toxicity of untested compounds.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning identifies phenotypic profile alterations of human dopaminergic neurons exposed to bisphenols and perfluoroalkyls 92%
- Discovery of Z1362873773: A Novel Fascin Inhibitor from a Large Chemical Library for Colorectal Cancer 91%
- Carcinogenicity and testicular toxicity of 2-bromopropane in a 26-week inhalation study using the rasH2 mouse model 91%
Similar papers in this journal
- Cell Painting and chemical structure read-across can complement each other for rat acute oral toxicity prediction in chemical early de-risking 95%
- A Method To Calibrate Chemical Agnostic Quantitative Adverse Outcome Pathways On Multiple Chemical Dose-Response Data 93%
- Cannabidiol Toxicity Driven by Hydroxyquinone Formation 92%
Similar papers in this journal
- Distinguishing classes of neuroactive drugs based on computational physicochemical properties and experimental phenotypic profiling in planarians 94%
- Exploring NCATS In-House Biomedical Data for Evidence-based Drug Repurposing 93%
- PharmaNet: Pharmaceutical discovery with deep recurrent neural networks. 91%
Similar papers in this journal
- Predicting Toxicity and Bioactivity of the Chemical Exposome: A Case Study for the Blood Exposome Database 96%
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 93%
- qHTSWaterfall: 3-dimensional visualization software for quantitative high-throughput screening (qHTS) data 91%
Similar papers in this journal
- NanoTox: Development of a parsimonious in silico model for toxicity assessment of metal-oxide nanoparticles using physicochemical features 94%
- Machine learning approaches identify chemical features for stage-specific antimalarial compounds 93%
- Support Vector Machine based prediction models for drug repurposing and designing novel drugs for colorectal cancer 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.