Back

Modeling nascent transcription from chromatin landscape and structure

Pielies Avelli, M.; Sigurdsson, A. I.; Narita, T.; Choudhary, C.; Rasmussen, S.

2024-06-06 bioinformatics
10.1101/2024.06.04.597340 bioRxiv
Show abstract

BackgroundDifferent cell types and their associated functionalities emerge from a single genomic sequence when certain regions are expressed while others remain silenced. Modeling gene expression and its potential malfunctioning in different cellular contexts is hence pivotal to understand both development and disease. ResultsWe present the Chromatin Landscape and Structure to Expression Regressor (CLASTER), an epigenetic-based deep neural network that can integrate different data modalities describing the chromatin landscape and its 3D structure. CLASTER effectively translates them into nascent transcription levels measured by EU-seq at a kilobasepair resolution. Our predictions reached a Pearson correlation with targets above r=0.86 at both bin and gene levels, without relying on DNA sequence nor explicitly extracted chromatin features. The model mostly used the information found within 10 kbp of the predicted locus, even when a wide genomic region of 1 Mbp was available. Explicit modeling of long-range interactions using multi-headed attention and high-resolution chromatin contact maps had little impact on model performance, despite the model correctly identifying elements in these inputs influencing nascent transcription. The trained model served then as a platform to predict the transcriptional impact of simulated epigenetic silencing perturbations. ConclusionsOur results point towards a rather local, integrative and combinatorial paradigm of gene regulation, where changes in the chromatin environment surrounding a gene shape its context-specific transcription. We conclude that the predominant locality and limitations of current machine learning approaches might emerge as a genuine signature of genomic organization, having broad implications for future modeling approaches.

Published in Genome Biology (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.