Back

Single-cell multi-scale footprinting reveals the modular organization of DNA regulatory elements

Hu, Y.; Ma, S.; Kartha, V. K.; Duarte, F. M.; Horlbeck, M.; Zhang, R.; Shrestha, R.; Labade, A.; Kletzien, H.; Meliki, A.; Castillo, A.; Durand, N.; Mattei, E.; Anderson, L. J.; Tay, T.; Earl, A. S.; Shoresh, N.; Epstein, C. B.; Wagers, A.; Buenrostro, J. D.

2023-03-29 genomics
10.1101/2023.03.28.533945 bioRxiv
Show abstract

Cis-regulatory elements control gene expression and are dynamic in their structure, reflecting changes to the composition of diverse effector proteins over time1-3. Here we sought to connect the structural changes at cis-regulatory elements to alterations in cellular fate and function. To do this we developed PRINT, a computational method that uses deep learning to correct sequence bias in chromatin accessibility data and identifies multi-scale footprints of DNA-protein interactions. We find that multi-scale footprints enable more accurate inference of TF and nucleosome binding. Using PRINT with single-cell multi-omics, we discover wide-spread changes to the structure and function of candidate cis-regulatory elements (cCREs) across hematopoiesis, wherein nucleosomes slide, expose DNA for TF binding, and promote gene expression. Activity segmentation using the co-variance across cell states identifies "sub-cCREs" as modular cCRE subunits of regulatory DNA. We apply this single-cell and PRINT approach to characterize the age-associated alterations to cCREs within hematopoietic stem cells (HSCs). Remarkably, we find a spectrum of aging alterations among HSCs corresponding to a global gain of sub-cCRE activity while preserving cCRE accessibility. Collectively, we reveal the functional importance of cCRE structure across cell states, highlighting changes to gene regulation at single-cell and single-base-pair resolution.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.