QuadStack: Specialized convolutional blocks enable in vivo BG4-binding motif prediction and highlight discrepancies with in vitro G-quadruplexes.
Ulas, P. N.; Doluca, O.
Show abstract
G-quadruplex (G4) prediction has been largely guided by in vitro biophysical rules, yet these models show limited agreement with in vivo measurements. Here, we present QuadStack, a deep learning model trained on a multi-study BG4-ChIP-seq compendium. QuadStack introduces two biologically grounded convolutional modules--G4Stack Convolution, which captures G/C stacking patterns, and Reverse Complement Convolution, which enforces strand-invariant representations consistent with ChIP-seq signals. QuadStack achieves strong predictive performance (AUC up to 0.94) and substantially outperforms widely used in vitro-based predictors on genomic test data. Beyond performance, our analyses reveal that BG4-associated sequence grammar is not solely governed by canonical isolated G-rich tracts, but also by patterns where G and C nucleotides are mixed. This suggests that cytosines are not simply disruptive in vivo, and raises the possibility that cytosines may play a context-dependent role or that guanines on the opposite strand contribute to the structure, which could explain the difference between in vivo and in vitro observations. Together these findings demonstrate a fundamental discrepancy between in vitro folding propensity and in vivo G4 biology, and establish QuadStack as both a predictive model and a framework for interpreting G4 formation in its native genomic context.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of upstream transcription factor binding sites in orthologous genes using mixed Student's t-test statistics 94%
- Base-resolution prediction of transcription factor binding signals by a deep learning framework 94%
- GRAFIMO: variant and haplotype aware motif scanning on pangenome graphs 94%
Similar papers in this journal
- PtWAVE: A High-Sensitive deconvolution software of sequencing trace for the Detection of Large Indels in Genome Editing 95%
- Rescuing Biologically Relevant Consensus Regions Across Replicated Samples 95%
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 95%
Similar papers in this journal
- Dissecting the binding mechanisms of transcription factors to DNA using a statistical thermodynamics framework. 95%
- LBFextract: unveiling transcription factor dynamics from liquid biopsy data 94%
- Modeling and analysis of site-specific mutations in cancer identifies known plus putative novel hotspots and bias due to contextual sequences 94%
Similar papers in this journal
- Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data 94%
- G4-iM Grinder: When size and frequency matter. G-Quadruplex, i-Motif and higher order structure search and analysis tool 94%
- A Comprehensive Evaluation of Self Attention for Detecting Regulatory Feature Interactions 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.