Iterative improvement of deep learning models using synthetic regulatory genomics
Ribeiro-dos-Santos, A. M.; Maurano, M. T.
Show abstract
Deep learning models can accurately reconstruct genome-wide epigenetic tracks from the reference genome sequence alone. But it is unclear what predictive power they have on sequence diverging from the reference, such as disease- and trait-associated variants or engineered sequences. Recent work has applied synthetic regulatory genomics to characterized dozens of deletions, inversions, and rearrangements of DNase I hypersensitive sites (DHSs). Here, we use the state-of-the-art model Enformer to predict DNA accessibility and RNA transcription across these engineered sequences when delivered at their endogenous loci. At high level, we observe a good correlation between accessibility predicted by Enformer and experimental data. But model performance was best for sequences that more resembled the reference, such as single deletions or combinations of multiple DHSs. Predictive power was poorer for rearrangements affecting DHS order or orientation. We use these data to fine-tune Enformer, yielding significant reduction in prediction error. We show that this fine-tuning retains strong predictive performance for other tracks. Our results show that current deep learning models perform poorly when presented with novel sequence diverging in certain critical features from their training set. Thus an iterative approach incorporating profiling of synthetic constructs can improve model generalizability and ultimately enable functional classification of regulatory variants identified by population studies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancers display constrained sequence flexibility and context-specific modulation of motif function 97%
- Identifying transcription factor-bound gene activators and silencers in the chromatin accessible human genome using ATAC-STARR-seq 96%
- Dissecting the regulatory activity and sequence content of loci with exceptional numbers of transcription factor associations 96%
Similar papers in this journal
- Enhlink infers distal and context-specific enhancer-promoter linkages 97%
- A global high-density chromatin interaction network reveals functional long-range and trans-chromosomal relationships 96%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.