Current genomic deep learning models display decreased performance in cell type specific accessible regions
Kathail, P.; Shuai, R. W.; Chung, R.; Ye, C. J.; Loeb, G.; Ioannidis, N. M.
Show abstract
AbstractO_ST_ABSBackgroundC_ST_ABSA number of deep learning models have been developed to predict epigenetic features such as chromatin accessibility from DNA sequence. Model evaluations commonly report performance genome-wide; however, cis regulatory elements (CREs), which play critical roles in gene regulation, make up only a small fraction of the genome. Furthermore, cell type specific CREs contain a large proportion of complex disease heritability. ResultsWe evaluate genomic deep learning models in chromatin accessibility regions with varying degrees of cell type specificity. We assess two modeling directions in the field: general purpose models trained across thousands of outputs (cell types and epigenetic marks), and models tailored to specific tissues and tasks. We find that the accuracy of genomic deep learning models, including two state-of-the-art general purpose models - Enformer and Sei - varies across the genome and is reduced in cell type specific accessible regions. Using accessibility models trained on cell types from specific tissues, we find that increasing model capacity to learn cell type specific regulatory syntax - through single-task learning or high capacity multi-task models - can improve performance in cell type specific accessible regions. We also observe that improving reference sequence predictions does not consistently improve variant effect predictions, indicating that novel strategies are needed to improve performance on variants. ConclusionsOur results provide a new perspective on the performance of genomic deep learning models, showing that performance varies across the genome and is particularly reduced in cell type specific accessible regions. We also identify strategies to maximize performance in cell type specific accessible regions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Connecting high-resolution 3D chromatin organization with epigenomics 96%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 96%
- Leveraging supervised learning for functionally-informed fine-mapping of cis-eQTLs identifies an additional 20,913 putative causal eQTLs 96%
Similar papers in this journal
- Epiphany: predicting Hi-C contact maps from 1D epigenomic signals 96%
- INFIMA leverages multi-omics model organism data to identify effector genes of human GWAS variants 96%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 96%
Similar papers in this journal
- Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation 97%
- Personal transcriptome variation is poorly explained by current genomic deep learning models 96%
- Benchmarking of deep neural networks for predicting personal gene expression from DNA sequence highlights shortcomings 96%
Similar papers in this journal
Similar papers in this journal
- Purifying selection on noncoding deletions of human regulatory elements detected using their cellular pleiotropy 95%
- scTIE: data integration and inference of gene regulation using single-cell temporal multimodal data 95%
- Quantitative occupancy of myriad transcription factors from one DNase experiment enables efficient comparisons across conditions 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.