Back

Identifying regulatory and spatial genomic architectural elements using cell type independent machine and deep learning models

Martens, L. D.; Faust, O.; Pirvan, L.; Bihary, D.; Samarajiwa, S. A.

2020-04-20 genomics
10.1101/2020.04.19.049585 bioRxiv
Show abstract

Chromosome conformation capture methods such as Hi-C enables mapping of genome-wide chromatin interactions and is a promising technology to understand the role of spatial chromatin organisation in gene regulation. However, the generation and analysis of these data sets at high resolutions remain technically challenging and costly. We developed a machine and deep learning approach to predict functionally important, highly interacting chromatin regions (HICR) and topologically associated domain (TAD) boundaries independent of Hi-C data in both normal physiological states and pathological conditions such as cancer. This approach utilises gradient boosted trees and convolutional neural networks trained on both Hi-C and histone modification epigenomic data from three different cell types. Given only epigenomic modification data these models are able to predict chromatin interactions and TAD boundaries with high accuracy. We demonstrate that our models are transferable across cell types, indicating that combinatorial histone mark signatures may be universal predictors for highly interacting chromatin regions and spatial chromatin architecture elements.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.