Back

Constrained Open Chromatin Regions Reveal Functional Elements Shaping Human Traits

Tomizuka, K.; Koido, M.; Suzuki, A.; Yoshino, S.; Tanaka, N.; Ishikawa, Y.; Liu, X.; Koyama, S.; Ishigaki, K.; Murakawa, Y.; Immune transcript/enhancer consortium (ITEC), ; Yamamoto, K.; Terao, C.

2025-05-16 genetics
10.1101/2024.10.31.621195 bioRxiv
Show abstract

Open chromatin regions (OCRs) define cell-type-specific regulatory elements across the genome, yet their functional significance varies, making it challenging to pinpoint biologically essential regions. Here, we introduce CAMBUS (Chromatin Accessibility Mutation Burden Score), a machine-learning framework that identifies active and evolutionarily constrained OCRs by leveraging surrounding DNA sequences. Applying CAMBUS to 29 immune cell types, we identified 66,043 constrained OCRs, which were substantially enriched in the known constraint genome (odds ratio=11.45 (95% confidence interval 9.33-14.05), P=4.7 x 10-68), while 90% of these OCRs were not prioritized by existing constraint metrics. These OCRs were highly enriched for known enhancers and super-enhancers, independent of known epigenetic markers and annotated regions, and overlapped with regulatory elements implicated in immune-mediated diseases and experimentally validated functional variants, including rare variants, particularly in leukocyte-related traits. Furthermore, CAMBUS revealed cell-type-specific transcriptional regulatory landscapes, linking genetic constraint with gene regulation in immune cells and identifying plausible connections between 1,533 causal variants and 70 complex traits. By defining biologically constrained regulatory elements at high resolution, CAMBUS provides a framework for understanding the selective pressures shaping the non-coding genome and its role in human health and disease.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.