Leveraging Targeted Gene Sets and Neural Networks for Zebrafish Transcriptome Extrapolation in High-Throughput Toxicogenomics
Howard, B. E.; Mav, D.; Balik-Meisner, M.; Phadke, D.; Green, A. J.; Truong, L.; Tanguay, R. L.; Shah, R. R.
Show abstract
BackgroundZebrafish (Danio rerio) are a powerful vertebrate model for developmental toxicology and chemical safety assessment, yet large-scale transcriptomics in zebrafish remains limited by cost and data heterogeneity. Targeted transcriptomics offers a cost-effective alternative, but gene extrapolation methods tailored to zebrafish have not been systematically developed or evaluated. ObjectivesWhile the S1500+ platform is widely used for toxicogenomics research with rat, mouse, and human cell lines as model systems, its use in zebrafish has been limited due to data scarcity and lack of suitable bioinformatics approaches for analysis of such data. To that end, we sought to (i) curate a large zebrafish transcriptomic training data resource, and (ii) evaluate multiple machine learning strategies for reconstructing unmeasured transcriptome-wide expression profiles for data originating from the zebrafish-specific reduced representation gene set ("Zf S1500+"). MethodsWe assembled 14,924 zebrafish RNA-Seq samples covering 21,930 genes across 1,246 studies. Using the Zf S1500+ gene subset (3,062 genes), we trained and tested three extrapolation approaches: principal components regression (PCR), a locally weighted extension of PCR (PCR+), and a neural network mixture-of-experts model (NN-MoE). Model performance was assessed using mean absolute error (MAE), mean squared regression error (MSRE), and weighted variants of these metrics. ResultsExtrapolation performance using the baseline approach was strongly influenced by tissue and developmental context, with within-tissue models outperforming cross-tissue models. Errors were lowest when training and testing were conducted within the same tissue or between developmentally related tissues. Both PCR+ and NN-MoE improved upon the baseline PCR approach, with NN-MoE reducing average MAE by [~]20% and MSRE by [~]17%. Importantly, extrapolation remained reliable for the majority of genes, even when limiting output to high-confidence predictions using an empirical MAE threshold. ConclusionsWe demonstrate that targeted transcriptomics can be effectively extended to zebrafish, enabling robust transcriptome-wide extrapolation at reduced cost. The NN-MoE method provided the most substantial gains, highlighting the value of non-linear and ensemble modeling in heterogeneous datasets. These results establish a scalable framework for zebrafish toxicogenomics and suggest that accuracy will continue to improve with larger, better-annotated datasets, paving the way for broader application in chemical safety assessments.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Defining the data gap: what do we know about environmental exposure, hazards and risks of pharmaceuticals in the European aquatic environment? 92%
- A Statistical Review of Virus Reduction in Coagulation, Flocculation, and Sedimentation Treatment Processes 90%
- Burden of Disease from Contaminated Drinking Water in Countries with High Access to Safely Managed Water: A Systematic Review 90%
Similar papers in this journal
- Deep autoencoder-based behavioral pattern recognition outperforms standard statistical methods in high-dimensional zebrafish studies 94%
- Longitudinal wastewater sampling in buildings reveals temporal dynamics of metabolites 90%
- Pathway analysis in metabolomics: pitfalls and best practice for the use of over-representation analysis 90%
Similar papers in this journal
- Picky with peakpicking: assessing chromatographic peak quality with simple metrics in metabolomics 90%
- Single sample pathway analysis in metabolomics: performance evaluation and application 89%
- Probabilistic quotient's work and pharmacokinetics' contribution: countering size effect in metabolic time series measurements 89%
Similar papers in this journal
- Network-based investigation of petroleum hydrocarbons-induced ecotoxicological effects and their risk assessment 94%
- Potential Systemic Availability Classification of Chemicals for Safety Assessment 93%
- Endocrine disrupting chemicals and COVID-19 relationships: a computational systems biology approach 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.