Contrastive Regulatory Embeddings Attention Model for Differential Expression Prediction with Personalized Genomes: Insights and Challenges
Hu, Z.; Ku, J.; Pollard, K.
Show abstract
Deep learning models applied to DNA sequences have achieved success in predicting gene expression, chromatin profiles, and variant pathogenicity. However, learning cross-individual differences remains challenging because DNA sequence variation between individuals is small, while gene expression is heavily influenced by non-genetic noise. Prior efforts to predict personalized gene expression from sequence have shown limited generalizability beyond training genes, revealing limitations such as the dilution of variant signals among highly similar input sequences across consecutive convolutional downsampling layers. In this work, we explore architectural modifications addressing these challenges. We propose a contrastive regulatory embedding attention model (CREAM) designed to better capture subtle sequence differences between individuals. To mitigate the impact of non-genetic variability, we decompose gene expression into genetic and non-genetic components and evaluate model performance on the genetic signal. Evaluated on simulated and GTEx transcriptomic datasets, CREAM captures tissue-specific gene expression and significantly outperforms baseline architectures on training genes. CREAM autonomously prioritizes statistically fine-mapped causal expression quantitative trait loci, and introducing an auxiliary L1 inductive bias further sharpens localizing causal variant in unseen genes. However, generalization to predicting expression of unseen genes collapses to near-zero correlation among all tested methods. Single-run model predictive uncertainty capture prediction accuracy in training genes and mirrors cross-run consistency in test genes. While accurate inference on unseen genes remains an open problem, our results highlight key obstacles and suggest directions for modeling personalized gene regulation from DNA sequence.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 96%
- Multi-scale deep tensor factorization learns a latent representation of the human epigenome 95%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.