Back

Deciphering the species-level structure of topologically associating domains

Singh, R.; Berger, B.

2021-10-29 bioinformatics
10.1101/2021.10.28.466333 bioRxiv
Show abstract

Chromosome conformation capture technologies such as Hi-C have revealed a rich hierarchical structure of chromatin, with topologically associating domains (TADs) as a key organizational unit, but experimentally reported TAD architectures, currently determined separately for each cell type, are lacking for many cell/tissue types. A solution to address this issue is to integrate existing epigenetic data across cells and tissue types to develop a species-level consensus map relating genes to TADs. Here, we introduce the TAD Map, a bag-of-genes representation that we use to infer, or "impute," TAD architectures for those cells/tissues with limited Hi-C experimental data. The TAD Map enables a systematic analysis of gene coexpression induced by chromatin structure. By overlaying transcriptional data from hundreds of bulk and single-cell assays onto the TAD Map, we assess gene coexpression in TADs and find that expressed genes cluster into fewer TADs than would be expected by chance, and show that time-course and RNA velocity studies further reveal this clustering to be strongest in the early stages of cell differentiation; it is also strong in tumor cells. We provide a probabilistic model to summarize any scRNA-seq transcriptome in terms of its TAD activation profile, which we term a TAD signature, and demonstrate its value for cell type inference, cell fate prediction, and multimodal synthesis. More broadly, our work indicates that the TAD Maps comprehensive, quantitative integration of chromatin structure and scRNA-seq data should play a key role in epigenetic and transcriptomic analyses. Software availability: https://tadmap.csail.mit.edu Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=187 HEIGHT=200 SRC="FIGDIR/small/466333v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1d652dborg.highwire.dtl.DTLVardef@1d9c68dorg.highwire.dtl.DTLVardef@7a8e66org.highwire.dtl.DTLVardef@1afaff_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.