Identifying regulatory and spatial genomic architectural elements using cell type independent machine and deep learning models
Martens, L. D.; Faust, O.; Pirvan, L.; Bihary, D.; Samarajiwa, S. A.
Show abstract
Chromosome conformation capture methods such as Hi-C enables mapping of genome-wide chromatin interactions and is a promising technology to understand the role of spatial chromatin organisation in gene regulation. However, the generation and analysis of these data sets at high resolutions remain technically challenging and costly. We developed a machine and deep learning approach to predict functionally important, highly interacting chromatin regions (HICR) and topologically associated domain (TAD) boundaries independent of Hi-C data in both normal physiological states and pathological conditions such as cancer. This approach utilises gradient boosted trees and convolutional neural networks trained on both Hi-C and histone modification epigenomic data from three different cell types. Given only epigenomic modification data these models are able to predict chromatin interactions and TAD boundaries with high accuracy. We demonstrate that our models are transferable across cell types, indicating that combinatorial histone mark signatures may be universal predictors for highly interacting chromatin regions and spatial chromatin architecture elements.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Towards Personalized Epigenomics: Learning Shared Chromatin Landscapes and Joint De-Noising of Histone Modification Assays 94%
- Genome wide clustering on integrated chromatin states and Micro-C contacts reveals chromatin interaction signatures 94%
- Prediction of G4 formation in live cells with epigenetic data: a deep learning approach 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.