Training deep learning models on personalized genomic sequences improves variant effect prediction
He, A. Y.; Palamuttam, N. P.; Danko, C. G.
Show abstract
Sequence-to-function models have broad applications in interpreting the molecular impact of genetic variation, yet have been criticized for poor performance in this task. Here we show that training models on functional genomic data with matched personal genomes improves their performance at variant effect prediction. Variant effect representations are retained even when fine tuning models to unseen cellular contexts and experimental readouts. Our results have implications for interpreting trait-associated genetic variation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HATCHet2: clone- and haplotype-specific copy number inference from bulk tumor sequencing data 96%
- Integrative epigenomic and functional characterization assay based annotation of regulatory activity across diverse human cell types 96%
- Multi-cell type deconvolution using a probabilistic model for single-molecule DNA methylation haplotypes 95%
Similar papers in this journal
Similar papers in this journal
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 96%
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 95%
- Inferring allele-specific copy number aberrations and tumor phylogeography from spatially resolved transcriptomics 95%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 95%
- Discovering single nucleotide variants and indels from bulk and single-cell ATAC-seq 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.