Back

ChemoCalib: multiblock PLS calibration of genome-scale metabolic models improves flux prediction over expression-only integration

zhang, X.

2026-07-31 bioinformatics
10.64898/2026.07.28.741216 bioRxiv
Show abstract

MotivationConstraint-based metabolic modeling faces a calibration gap: genome-scale metabolic models (GEMs) integrated with transcriptomics alone rely on expression-to-flux heuristics (E-Flux, GIMME, iMAT, MOMENT) that ignore cross-omics co-variance structure and lack statistical mechanisms for propagating omics uncertainty into reaction bounds, yielding flux predictions with limited agreement to 13C metabolic flux analysis (MFA) measurements. ResultsWe present Chemo-Calib, a multiblock PLS (MB-PLS) framework that calibrates GEM reaction bounds from the shared latent structure of metabolomics, transcriptomics, and proteomics data. On 11 E. coli 13C-MFA reference conditions spanning the Keio fluxome and Holm 2010 datasets, ChemoCalib constrained FBA on iJO1366 achieves a Spearman{rho} = 0.461 overall (up to 0.523 in PPP) and Pearson r of 0.49-0.58 across central carbon pathways, with statistically significant improvement over expression-only baselines including E-Flux2 and SPOT (p < 0.05, Holm-corrected). The latent-to-constraint mapping employs GPR-aware VIP aggregation (Algorithm 1) to project multi-omics latent scores onto genome-scale reaction bounds without heuristic thresholding. An optional in-silico active learning loop (relegated to Supplementary Material) further tightens calibration through virtual experiment selection. AvailabilityChemoCalib is open-source (MIT) at https://github.com/chemocalib/chemocalib with Docker support, a 5-minute tutorial, and pre-computed iJO1366 benchmarks. Preprint available at bioRxiv; code archived at Zenodo DOI: 10.5281/zenodo.21645890.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.