Back

Host-aware Identification of Intrinsic Gene Expression Biopart Parameters from Combinatorial Libraries

Pico, J.; Arboleda-Garcia, A.; Penas, D. R.; Banga, J. R.; Vignoni, A.; Boada, Y.

2026-01-21 synthetic biology
10.64898/2026.01.21.700808 bioRxiv
Show abstract

Model-based design in synthetic biology is limited by the lack of quantitative, mechanistically interpretable biopart parameters that remain valid across genetic and physiological contexts. This limitation is particularly acute for transcriptional units, whose expression phenotypes emerge from nonlinear coupling between plasmid copy number, transcription, translation, and host resource allocation. Here we introduce a context- and host-aware framework for the absolute characterisation of gene expression bioparts embedded in combinatorial libraries of constitutive transcriptional units. Our approach leverages a digital twin of Escherichia coli conditioned on experimentally measured growth rate, used as a low-dimensional physiological proxy for cellular state. By embedding this growth-conditioned digital twin in a model-in-the-loop identification strategy, host-circuit interactions are explicitly accounted for and decoupled from intrinsic biopart properties. Using structured combinatorial libraries, we identify biophysically interpretable and transferable parameters for plasmid origins, promoters, and ribosome binding sites. In particular, we uncover an intrinsic translation initiation capacity of ribosome binding sites that remains invariant across genetic contexts and growth conditions, while context-dependent translation rates emerge as physiological projections of this invariant descriptor. This intrinsic parameterisation enables accurate prediction of protein synthesis across diverse host states, supports incremental and patchwork library expansion, and reveals localized failures of modularity that are obscured by phenotype-only characterisation. Together, these results establish a principled link between DNA sequence, intrinsic biopart parameters, and circuit-level phenotypes, providing a scalable and host-aware foundation for predictive design in synthetic biology.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.