Back

PREDICT: Advancing Accurate Gene Expression Prediction and Motif Identification in Plant Stress Responses

Wu, T.-Y.; Liu, M.-J.; Thalimaraw, L.; Eo, W. X. H.

2024-03-31 bioinformatics
10.1101/2024.03.28.587275 bioRxiv
Show abstract

Cells respond to environmental stimuli through transcriptional responses, orchestrated by transcription factors (TFs) that interpret the gene cis-regulatory DNA sequences, determining gene expression dynamics timing and locations. Diversification in TFs and cis-regulatory element (CRE) interactions result in unique gene regulatory networks (GRNs) that underpin plant adaptation. A primary challenge is identifying Transcription Factor Binding Motifs (TFBMs) for temporal and condition-specific gene expressions in plants. While the Multiple EM for Motif Elicitation (MEME) suite identifies stress-responsive CREs in Arabidopsis, its predictive power for gene expression remains uncertain. Alternatively, the k-mer approach identifies CRE sites and consensus TF motifs, thereby improving gene expression prediction models. In this study, we harnessed the power of a k-mer pipeline to address sequence-to-expression prediction problems across diverse abiotic stresses, in both bryophytic and vascular plants, including monocots and dicots. Moreover, we characterized both un-gapped and gapped CREs and, coupled with GRN analyses, pinpointed key TFs within transcriptional cascades. Lastly, we developed the Predictive Regulatory Element Database for Identifying Cis-regulatory elements and Transcription factors (PREDICT), a web tool for efficient k-mer identification. This advancement will enrich our understanding of the cis-regulatory code landscape that shapes gene regulation in plant adaptation. PREDICT web tool is available at [http://predict.southerngenomics.org/kmers/kmers.php].

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.