Pretraining Improves Prediction of Genomic Datasets Across Species
Huang, F.; Wang, Y.; Song, J. H.; Cutkosky, A.
Show abstract
Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this stream-lined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this compromised performance could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cross-species imputation and comparison of single-cell transcriptomic profiles 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
- Enhancement of network architecture alignment in comparative single-cell studies 96%
Similar papers in this journal
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 96%
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 96%
- Towards Universal Cell Embeddings: Integrating Single-cell RNA-seq Datasets across Species with SATURN 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.