accuEnhancer: Accurate enhancer prediction by integration of multiple cell type data with deep learning
Tung, Y.-A.; Yang, W.-T.; Hsieh, T.-T.; Chang, Y.-C.; Wu, J.-T.; Oyang, Y.-J.; Chen, C.-Y.
Show abstract
Enhancers are one class of the regulatory elements that have been shown to act as key components to assist promoters in modulating the gene expression in living cells. At present, the number of enhancers as well as their activities in different cell types are still largely unclear. Previous studies have shown that enhancer activities are associated with various functional data, such as histone modifications, sequence motifs, and chromatin accessibilities. In this study, we utilized DNase data to build a deep learning model for predicting the H3K27ac peaks as the active enhancers in a target cell type. We propose joint training of multiple cell types to boost the model performance in predicting the enhancer activities of an unstudied cell type. The results demonstrated that by incorporating more datasets across different cell types, the complex regulatory patterns could be captured by deep learning models and the prediction accuracy can be largely improved. The analyses conducted in this study demonstrated that the cell type-specific enhancer activity can be predicted by joint learning of multiple cell type data using only DNase data and the primitive sequences as the input features. This reveals the importance of cross-cell type learning, and the constructed model can be applied to investigate potential active enhancers of a novel cell type which does not have the H3K27ac modification data yet. AvailabilityThe accuEnhancer package can be freely accessed at: https://github.com/callsobing/accuEnhancer
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Capturing large genomic contexts for accurately predicting enhancer-promoter interactions 95%
- DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding 95%
- PTFSpot: Deep co-learning on transcription factors and their binding regions attains impeccable universality in plants 94%
Similar papers in this journal
- Systematic Prediction of Regulatory Motifs from Human ChIP-Sequencing Data Based on a Deep Learning Framework 96%
- Assessing base-resolution DNA mechanics on the genome scale 95%
- ANANSE: An enhancer network-based computational approach for predicting key transcription factors in cell fate determination 94%
Similar papers in this journal
- TIVAN-indel: A computational framework for annotating and predicting noncoding regulatory small insertion and deletion 95%
- DeepPHiC: Predicting promoter-centered chromatin interactions using a novel deep learning approach 95%
- Controlled Noise: Evidence of Epigenetic Regulation of Single-Cell Expression Variability 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.