Reconstructing sequence-grammar trajectories enables interpretable and tunable cis- regulatory element design
Ma, M.; Bu, W.; Liu, G.; Liu, Y.; Liu, S.; Zhao, Z.; Yao, S.; Hua, Q.; Zhang, Y.; Zhong, C.; Huang, H.; Deng, P.; Jin, P.; Yin, Q.; Cao, C.; Liu, H.; Xu, M.; He, Y.; Qin, T.; Chen, Z.
Show abstract
Designing synthetic cis-regulatory elements (CREs) with cell-type-specific activity remains challenging, and optimization is usually treated as a black box, obscuring how regulatory grammar emerges and why design trajectories fail. Here, we present GO-CRE (Guided Optimization of Cis-Regulatory Elements), which combines efficient sequence generation, predictor-guided reinforcement learning, and trajectory-level interpretation. GO-CRE reconstructs iterative sequence changes in a shared sequence-grammar landscape and identifies coordinated update programs corresponding to search, commitment, and optimization. Productive trajectories in HepG2 and K562 progressively acquired cell-type-associated grammar, whereas SK-N-SH trajectories remained confined to local basins. In HepG2, trajectory analysis revealed a low-complexity polyG trap; introducing a polyG penalty redirected optimization toward HNF/FOXA-associated features. Final designs retained sequence diversity while converging on cell-type-associated motif patterns. Lentiviral MPRA validated cell-type-specific activity in K562 and HepG2 and showed higher average activity of HepG2 designs than endogenous CREs. Together, these findings establish sequence-grammar trajectory reconstruction as a basis for interpretable and tunable synthetic CRE design.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals 96%
- Large-scale DNA-based phenotypic recording and deep learning enable highly accurate sequence-function mapping 96%
- Structure-guided engineering of type I-F CASTs for targeted gene insertion in human cells 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.