Back

Hybrid genome-scale modeling and machine learning reveal cost-efficient strategies for phototrophic PHB production in Rhodopseudomonas palustris

Hernandez Gonzalez, H. A.; Buitron, G.

2026-04-26 biochemistry
10.64898/2026.04.24.720730 bioRxiv
Show abstract

Plastic pollution from fossil-based materials is a major global environmental challenge. Microbial-derived bioplastics, such as polyhydroxybutyrate (PHB), offer a promising biodegradable alternative. However, the high substrate and operational costs of PHB production remain a major barrier to large-scale deployment. Optimizing PHB synthesis requires navigating a multidimensional design space of metabolic, nutritional, and operational variables that is impractical to explore experimentally. Here, we developed an integrated computational framework that combines a genome-scale metabolic model (GEM) of Rhodopseudomonas palustris, machine-learning surrogate modeling, Pareto multi-objective optimization, and thermodynamics-based flux analysis (TFA) to identify cost-efficient and biologically feasible PHB production strategies. Experimental and literature-derived medium compositions were translated into mechanistic constraints, enabling the GEM to generate metabolically coherent synthetic datasets that augmented sparse experimental observations. CatBoost surrogate models trained on this hybrid dataset accurately predicted PHB synthesis across thousands of hypothetical conditions, and Pareto optimization revealed operating regimes that balance PHB productivity with nutrient cost. TFA validated the thermodynamic feasibility of these strategies and refined pathway usage, reinforcing thiolase-initiated routing into PHB biosynthesis and suppressing infeasible {beta}-oxidation-like redox loops. Overall, this hybrid GEM-ML-TFA framework identifies metabolic bottlenecks, engineering targets, and cost-optimal nutrient regimes for phototrophic PHB production, providing a scalable blueprint for rational process and strain design.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
ACS Synthetic Biology
287 papers in training set
Top 0.4%
9.9%
2
Nature Communications
5641 papers in training set
Top 19%
9.0%
3
Water Research
79 papers in training set
Top 0.2%
8.0%
4
Applied and Environmental Microbiology
339 papers in training set
Top 1%
6.8%
5
mSystems
394 papers in training set
Top 1%
5.6%
6
Bioresource Technology
12 papers in training set
Top 0.1%
5.6%
7
Microbial Cell Factories
27 papers in training set
Top 0.1%
4.4%
8
PLOS ONE
5266 papers in training set
Top 42%
2.5%
50% of probability mass above
9
Scientific Reports
3612 papers in training set
Top 43%
2.4%
10
New Phytologist
346 papers in training set
Top 3%
2.2%
11
Metabolic Engineering
75 papers in training set
Top 0.4%
2.2%
12
ACS Catalysis
18 papers in training set
Top 0.1%
2.2%
13
Environmental Science & Technology
64 papers in training set
Top 0.6%
1.9%
14
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 27%
1.8%
15
Biotechnology and Bioengineering
53 papers in training set
Top 0.4%
1.8%
16
ACS Chemical Biology
167 papers in training set
Top 2%
1.5%
17
Microbial Biotechnology
34 papers in training set
Top 0.5%
1.5%
18
Biochemistry
148 papers in training set
Top 2%
1.4%
19
Nature Chemical Biology
119 papers in training set
Top 2%
1.1%
20
Communications Chemistry
48 papers in training set
Top 1%
1.1%
21
Journal of Biological Chemistry
690 papers in training set
Top 8%
1.1%
22
Science Advances
1243 papers in training set
Top 28%
1.0%
23
The ISME Journal
228 papers in training set
Top 3%
1.0%
24
Synthetic and Systems Biotechnology
11 papers in training set
Top 0.2%
0.9%
25
Algal Research
21 papers in training set
Top 0.3%
0.9%
26
PLOS Computational Biology
1863 papers in training set
Top 20%
0.9%
27
Communications Biology
993 papers in training set
Top 29%
0.9%
28
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 2%
0.9%
29
mBio
833 papers in training set
Top 11%
0.9%
30
Nature Chemistry
42 papers in training set
Top 1%
0.6%