Back

A database of over 15.000 strain design publications reveals a conserved set of metabolic engineering targets across microbial hosts and products

Marquez-Zavala, E.; Di Bartolomeu, F.; Machado, D.

2025-12-16 systems biology
10.64898/2025.12.15.694291 bioRxiv
Show abstract

Microbial biotechnology has the potential to address several societal issues through the sustainable production of industrially relevant compounds. Despite decades of successful cases, rational engineering of microbial metabolism is still a complex process due to the fine balance between nutrient supply, allocation of cellular resources, energy demand and redox balancing. In this work, we implemented a text-mining workflow for metabolic engineering and compiled a database of experimentally validated strain design strategies from over 15.000 research articles, which includes information on host strain, target compounds, and gene modifications. This large dataset reveals trends on the selection of suitable hosts for different kinds of products and the respective gene targets. Despite the wide variety of microbes and products, we observe a conserved set of target metabolic genes associated with central carbon metabolism, especially in upper glycolysis, pentose-phosphate pathway, citric acid cycle, and fermentative pathways. The most distinguishing feature among strain design strategies seems not to be which genes are targeted, but rather the direction in which they are modified (increased or decreased expression). Controlling flux at key branching points and redox balancing reactions is thus a critical engineering step to steer metabolism. Our collection of 25 years of literature can provide a stepping stone for starting new strain design projects without reinventing the wheel.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.