Gene-centered identification of cis-regulatory islands reveals regulatory landscapes complementary to motif-centric approaches
Ohmori, Y.
Show abstract
Cis-regulatory elements constitute a fundamental layer of gene regulation, yet their computational identification has largely relied on transcription factor (TF)-centric frameworks that assume genome-wide background normalization and explicit TF binding models. While effective at the genome scale, such assumptions are less appropriate for gene-centered analyses, where local sequence composition rather than global averages defines the relevant regulatory context. Here, we introduce a TF-independent framework for the gene-centered identification of cis-regulatory islands (GCIC), which detects regulatory structure based on the local enrichment and diversity of short cis-regulatory sequence words derived from curated plant regulatory elements. Cis-regulatory islands are identified through the spatial overlap of independently enriched motif families, without relying on TF identity, binding affinity, or genome-wide normalization. Application of the GCIC framework to the DROOPING LEAF (DL) locus in rice identifies discrete cis-regulatory islands, including one that coincides with a previously characterized intronic regulatory region, and reveals spatial patterns distinct from those detected by PWM-based motif scanning and motif clustering approaches. Genome-wide analyses further show that cis-regulatory islands are broadly distributed across genes but exhibit heterogeneous motif-family usage: regulatory vocabulary diversity expands at the gene level, whereas individual islands preferentially reuse a limited set of motif-family combinations. These results indicate that cis-regulatory organization is best described as a gene-centered property of sequence vocabulary usage, in which regulatory diversity arises through gene-specific deployment and constrained reuse of motif-family combinations rather than unrestricted combinatorial complexity. The GCIC framework thus provides a complementary representation of regulatory landscapes tailored to gene-centered analyses, capturing regulatory features that are not readily detected by motif-centric approaches optimized for genome-wide inference.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Diverse patterns of secondary structure across genes and transposable elements are associated with siRNA production and epigenetic fate 94%
- Genome-wide nucleosome and transcription factor responses to genetic perturbations reveal chromatin-mediated mechanisms of transcriptional regulation 93%
- Simple and efficient measurement of transcription initiation and transcript levels with STRIPE-seq 93%
Similar papers in this journal
- Functional identification of cis-regulatory long noncoding RNAs at controlled false-discovery rates 93%
- Identification of transcription factor co-binding patterns with non-negative matrix factorization 93%
- Best practices for perturbation MPRA--a computational evaluation framework of sequence design strategies 92%
Similar papers in this journal
- The cis-regulatory codes of response to combined heat and drought stress in Arabidopsis thaliana 94%
- Accurate prediction of cis-regulatory modules reveals a prevalent regulatory genome of humans 94%
- ChEC-seq2: an improved Chromatin Endogenous Cleavage sequencing method and bioinformatic analysis pipeline for mapping in vivo protein-DNA interactions 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.