Billion-Scale Deciphering of Human Gene Regulatory Grammar
Mayne, J.; Exposito Rodriguez, M.; Burrell, M.; Ceroni, F.; Goldrick, S.
Show abstract
Predicting how DNA sequence specifies gene expression remains a core challenge across regulatory genomics. Most predictive assays and models depend on native genomic DNA, constraining the full biochemical engineering space for assessing and designing new sequences. Here, we address this gap with a scalable experimental-computational platform that rapidly generates million-scale sequence-to-expression datasets that directly link degenerate sequences to their function in human cells. We built degenerate libraries of 200-bp promoter cassettes and performed pooled stable integration of up to 1012 unique constructs, enabling the curation of million-scale sequence-to-expression datasets by fluorescently sorting billions of human cells. Biophysical modeling of transcription-factor occupancy on the data using position weight matrices reveals a broad spectrum of correlations between factor abundance and expression levels, with some co-abundances reaching Pearsons r {approx} 0.99, consistent with cooperative and probabilistic regulation. Leveraging the dataset, we trained sequence-to-expression deep learning models that predict held-out expression with Pearson r {approx} 0.4, converge on shared sequence determinants, and agree strongly with each other (Pearsons r = 0.93), indicating reproducible sequence-expression relationships. Finally, with minimal retraining the models generalize to an independently generated dataset collected under distinct sorting conditions, transferring sequence rules across contexts. Our platform enables repeated, rapid studies and supports deeper mechanistic insight while providing baseline models for forward design of human regulatory elements, advancing prediction beyond genomic-DNA-anchored methods.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- An endoribonuclease-based feedforward controller for decoupling resource-limited genetic modules in mammalian cells 96%
- Systematic analysis of low-affinity transcription factor binding site clusters in vitro and in vivo establishes their functional relevance 96%
- Model-driven generation of artificial yeast promoters 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.