Back

Prediction-Guided Design of a More Developable FGF21 Construct

Bozkurt, C.; Nathanail, E.; Goteti, A.

2026-07-14 bioengineering
10.64898/2026.07.13.738140 bioRxiv
Show abstract

For structural-biology and protein-production pipelines, the hardest part of a difficult protein is not the biology -- it is obtaining a well-behaved sample for functional studies. Programs routinely stall at construct design, expression, and purification: deciding where to truncate, which tags to use, how to express, and how to purify so the protein survives concentration and handling. These decisions are still made largely by literature precedent and experimental experience, and they require trial-and-error before arriving at a functional construct for hard targets. We present a prospective, single-pair wet-lab case study testing whether an integrated computational platform can improve these decisions. For human fibroblast growth factor 21 (FGF21) -- a clinically important and stability-challenged metabolic hormone -- we compared two expression constructs produced side by side under the same experimental workflow, using two different design strategies: one designed by a scientist from the literature (reproducing the published core-domain construct, PDB 6M6E), and one designed by the Orbion platform -- an AI, prediction-guided protein-design system (orbion.life) -- which additionally generated the expression and purification protocols (executed scientist-in-the-loop). The platforms construct used an unconventional, longer C-terminal boundary not found in public sequence databases. Since the two constructs differ in more than one feature, we treat them as workflow-level designs throughout. The scientist construct gave a higher initial yield ([~]2.4 xmore protein recovered at affinity capture). The platform-designed construct, however, showed a more favourable downstream developability profile: it concentrated higher (1.4 vs 0.7 mg/mL) while remaining more monodisperse by dynamic light scattering (DLS). The scientist construct, in contrast, aggregated on concentration, so its initial-yield advantage did not survive: in the final concentrated sample the Orbion construct provided the more usable material for downstream studies. Computed for the mammalian host used, the platform had prospectively scored its own design higher (composite 68.7 vs 59.0 for the scientist-designed construct), and its predictions of yield, solubility, and disorder matched the wet-lab outcome. This is a single, deliberately scoped case study, not a population-level benchmark; the two constructs differ in more than one feature, and biological activity was not assayed. Alongside the bottlenecks of this approach discussed here, used as a decision aid, prediction-guided construct and protocol design has the potential to remove costly iteration cycles of protein production campaigns.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Protein Science
246 papers in training set
Top 0.1%
31.6%
2
mAbs
32 papers in training set
Top 0.1%
6.8%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.4%
5.6%
4
PLOS ONE
5266 papers in training set
Top 30%
5.2%
5
Scientific Reports
3612 papers in training set
Top 29%
3.6%
50% of probability mass above
6
ACS Synthetic Biology
287 papers in training set
Top 0.9%
3.3%
7
Methods
34 papers in training set
Top 0.1%
3.2%
8
Biotechnology and Bioengineering
53 papers in training set
Top 0.3%
2.4%
9
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 0.9%
1.9%
10
ACS Omega
105 papers in training set
Top 1%
1.9%
11
Trends in Biotechnology
12 papers in training set
Top 0.1%
1.9%
12
Journal of Molecular Biology
232 papers in training set
Top 2%
1.8%
13
PeerJ
308 papers in training set
Top 5%
1.8%
14
Communications Chemistry
48 papers in training set
Top 0.5%
1.8%
15
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
1.8%
16
New Biotechnology
12 papers in training set
Top 0.1%
1.5%
17
Synthetic Biology
24 papers in training set
Top 0.2%
1.1%
18
International Journal of Molecular Sciences
494 papers in training set
Top 11%
1.1%
19
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
20
International Journal of Biological Macromolecules
76 papers in training set
Top 2%
1.1%
21
Nature Communications
5641 papers in training set
Top 56%
0.9%
22
Communications Biology
993 papers in training set
Top 29%
0.9%
23
Bioinformatics
1204 papers in training set
Top 9%
0.6%
24
iScience
1154 papers in training set
Top 38%
0.6%