Back

Billion-Scale Deciphering of Human Gene Regulatory Grammar

Mayne, J.; Exposito Rodriguez, M.; Burrell, M.; Ceroni, F.; Goldrick, S.

2025-11-10 synthetic biology
10.1101/2025.11.10.687627 bioRxiv
Show abstract

Predicting how DNA sequence specifies gene expression remains a core challenge across regulatory genomics. Most predictive assays and models depend on native genomic DNA, constraining the full biochemical engineering space for assessing and designing new sequences. Here, we address this gap with a scalable experimental-computational platform that rapidly generates million-scale sequence-to-expression datasets that directly link degenerate sequences to their function in human cells. We built degenerate libraries of 200-bp promoter cassettes and performed pooled stable integration of up to 1012 unique constructs, enabling the curation of million-scale sequence-to-expression datasets by fluorescently sorting billions of human cells. Biophysical modeling of transcription-factor occupancy on the data using position weight matrices reveals a broad spectrum of correlations between factor abundance and expression levels, with some co-abundances reaching Pearsons r {approx} 0.99, consistent with cooperative and probabilistic regulation. Leveraging the dataset, we trained sequence-to-expression deep learning models that predict held-out expression with Pearson r {approx} 0.4, converge on shared sequence determinants, and agree strongly with each other (Pearsons r = 0.93), indicating reproducible sequence-expression relationships. Finally, with minimal retraining the models generalize to an independently generated dataset collected under distinct sorting conditions, transferring sequence rules across contexts. Our platform enables repeated, rapid studies and supports deeper mechanistic insight while providing baseline models for forward design of human regulatory elements, advancing prediction beyond genomic-DNA-anchored methods.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.