Back

Comp2GPR: A Sequence-Driven Framework for Gene.Protein-Reaction Rule Reconstruction

Castillo, S.

2026-06-26 bioinformatics
10.64898/2026.06.24.734174 bioRxiv
Show abstract

Accurate gene-protein-reaction (GPR) associations are essential for the predictive performance of genome-scale metabolic models (GEMs),as they define the mapping between genes, enzymes, and metabolic reactions. However, GPR rules are often incomplete or inconsistent due to limitations in annotation transfer and the ambiguous representation of multi-subunit protein complexes, leading to errors in downstream analyses such as gene essentiality prediction. Here, I introduce Comp2GPR, an automated pipeline for reconstructing GPR rules that integrates curated protein complex information with sequence-level evidence. Protein complexes were sourced from the Complex Portal and subjected to an AI-assisted curation workflow to retain only metabolically relevant assemblies. Comp2GPR combines deterministic sequence similarity mapping with explicit rule construction to generate Boolean GPR expressions that accurately represent obligate subunit relationships and isoenzyme redundancy. I evaluated the impact of the reconstructed GPR rules by integrating them into the Yeast9 metabolic model and comparing gene essentiality predictions with the original model. While global performance metrics remained largely unchanged, the updated model achieved a net improvement in prediction accuracy through gene-level corrections. Overall, Comp2GPR demonstrates that combining curated protein complex data with sequence-based validation improves the accuracy, interpretability, and reproducibility of GPR rules. The method provides a robust framework for enhancing metabolic model annotations and supports more reliable simulation-based analyses.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 0.6%
26.6%
2
Bioinformatics
1204 papers in training set
Top 2%
11.9%
3
Bioinformatics Advances
203 papers in training set
Top 0.3%
8.9%
4
BMC Bioinformatics
457 papers in training set
Top 1%
7.3%
50% of probability mass above
5
Nucleic Acids Research
1281 papers in training set
Top 5%
4.0%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.8%
7
Nature Communications
5641 papers in training set
Top 39%
2.4%
8
mSystems
394 papers in training set
Top 3%
2.4%
9
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.1%
10
Cell Systems
201 papers in training set
Top 2%
2.1%
11
PLOS ONE
5266 papers in training set
Top 47%
1.9%
12
Molecular Systems Biology
162 papers in training set
Top 1%
1.7%
13
Briefings in Bioinformatics
354 papers in training set
Top 4%
1.7%
14
Metabolic Engineering
75 papers in training set
Top 0.5%
1.5%
15
Genome Biology
637 papers in training set
Top 7%
1.1%
16
ACS Synthetic Biology
287 papers in training set
Top 2%
1.1%
17
Scientific Reports
3612 papers in training set
Top 65%
1.1%
18
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
19
Molecular Biology of the Cell
311 papers in training set
Top 3%
1.1%
20
npj Systems Biology and Applications
125 papers in training set
Top 2%
1.0%
21
G3: Genes, Genomes, Genetics
252 papers in training set
Top 4%
0.8%
22
iScience
1154 papers in training set
Top 35%
0.8%
23
GENETICS
483 papers in training set
Top 4%
0.8%
24
in silico Plants
27 papers in training set
Top 0.3%
0.6%
25
Journal of Chemical Information and Modeling
238 papers in training set
Top 3%
0.6%
26
Genome Research
468 papers in training set
Top 7%
0.6%
27
Cell Reports Methods
165 papers in training set
Top 5%
0.6%