Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data
Rastogi, R.; Reddy, A. J.; Chung, R.; Ioannidis, N. M.
Show abstract
Genomic sequence-to-expression deep learning models, which are trained to predict gene expression and other molecular phenotypes across the reference genome, have recently been shown to have poor out-of-the-box performance in predicting gene expression variation across individuals based on their personal genome sequences. Here we explore whether additional training (fine-tuning) on paired personal genome and transcriptome data improves the performance of such sequence-to-expression models. Using Enformer as a representative pretrained model, we explore various fine-tuning strategies. Our results show that fine-tuning improves cross-individual prediction performance over the baseline Enformer model for held-out individuals on genes seen during fine-tuning, with comparable performance to variant-based linear models commonly used in transcriptome-wide association studies. However, fine-tuning does not improve model generalizability on held-out genes, which contain sequences and variants unseen during fine-tuning, highlighting a remaining open challenge in the field.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scalable unified framework of total and allele-specific counts for cis-QTL, fine-mapping, and prediction 97%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 97%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.