Natural variation in regulatory code revealed through Bayesian analysis of plant pan-genomes and pan-transcriptomes
Wei, W.; Wu, X.; Sutherland, C. A.; Lin, Y.; Lunde, C.; Exposito-Alonso, M.; Krasileva, K.
Show abstract
Understanding the genetic code of cis-regulatory elements (CREs) is essential for engineering gene expression and modulating agronomic traits in crops. In plants, CREs underlying rapid evolution of gene expression often overlap with structural variation in promoters, making them undetectable using single-reference genomes. Here, we develop K-PROB (K-mer-based in silico PROmoter Bashing), a computational tool that learns from intraspecies promoter sequence and gene expression variation in pan-genomes and pan-transcriptomes to identify CREs controlling gene expression. K-PROB deploys a k-mer-based Bayesian variable selection framework to prioritize causal variable identification. We demonstrate the effectiveness of our approach in maize and soybean, two staple crops species. Applying K-PROB to genes with the most highly variable promoter sequences and the most diverse patterns of expression, such as nucleotide-binding leucine-rich repeat receptors, we identified k-mers enriched for bona fide transcription factor binding sequences, and overlapping with open chromatin regions and DAP-seq binding sites. Notably, multiple significant k-mers are located within presence/absence structural variants, highlighting structural variation in promoters as key drivers of transcriptional diversity of highly variable genes. We further validated the regulatory effects of identified k-mers on gene expression using luciferase reporter assays. Our results showcase a high-throughput and pangenomic approach for probing natural intraspecies cis-regulatory diversity, discovering new causative cis-elements, and facilitating future expression engineering across plant species. Significance StatementUnderstanding which DNA sequences control gene expression is essential for crop improvement. Current methods for identifying regulatory elements rely on expensive, specialized biochemical datasets typically limited to a single genotype. We developed a computational tool that links natural sequence variation and gene expression variation to identify functional regulatory sequences. Our tool employs a statistical framework that prioritizes causality over correlation, in contrast to most genome-wide association studies. Applying it to maize and soybean, two staple crops, we uncovered known and novel regulatory elements and validated them with molecular assays. Our approach is scalable, cost-effective, and efficiently utilizes natural variation from existing pangenomic datasets, opening new avenues for future crop engineering and studying gene regulation in diverse plant species.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The cis-regulatory codes of response to combined heat and drought stress in Arabidopsis thaliana 97%
- DiffSegR: An RNA-Seq data driven method for differential expression analysis using changepoint detection 94%
- Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs 94%
Similar papers in this journal
- Genome-wide, Organ-delimited gene regulatory networks (OD-GRNs) provide high accuracy in candidate TF selection across diverse processes. 96%
- Premeiotic 24-nt phasiRNAs are present in the Zea genus and unique in biogenesis mechanism and molecular function 96%
- Genome-Wide Profiling of Soybean WRINKLED1 Transcription Factor Binding Sites Provides Insight into Seed Storage Lipid Biosynthesis 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.