Back

Geometric Theoretical Framework for Dynamic Protein Mutation Detection Models: Defect Awareness and Pathogenicity Prediction

Shao, H.

2026-04-26 bioinformatics
10.64898/2026.04.22.720255 bioRxiv
Show abstract

Traditional protein mutation detection and pathogenicity prediction pipelines rely on static single-conformation structural modeling, inherently ignoring conformational flexibility, dynamic ensemble evolution, and the underlying manifold geometry of protein dynamics. This induces systematic detection failures in flexible regions, allosteric sites, and metastable functional domains, yet lacks a rigorous mathematical characterization of such failure mechanisms. In this work, we establish a theorem-driven geometric-algebraic framework for dynamic protein mutation modeling. Starting from a dynamic conformational Riemannian manifold, we construct the latent representation space via representation-induced completion of operator-valued observations, rather than pre-assumed embedding structures. Within this setting, algebraic constraints are not imposed axiomatically but relaxed into learnable approximate Lie algebra regularization, enabling statistical verification of structural consistency. By integrating Levi-Civita connection, geodesic deviation, and heat kernel asymptotics, we introduce a Lipschitz-stable topological spectral defect (TSD,{delta} spec) index that quantifies the intrinsic inconsistency between static representations and dynamic geometric invariants, linking it to curvature-induced instability and Lie algebra deformation. Under a functorial compatibility principle, we design a dual-branch architecture for pathogenicity prediction and defect awareness, realized via local Lie algebra encoding and low-rank spectral approximation. On multi-source datasets (108 curated PDB structures, 1060 validated residues from ClinVar, DMS, MaveDB, and gnomAD), we establish three fundamental theorems and validate key findings: TSD effectively distinguishes pathogenic/functional variants (PTM: {micro} = 0.386, OMIM: {micro} = 0.443, Clin-Var: {micro} = 0.302) from neutral ones (gnomAD: {micro} = -0.660) with high significance (P = 6.67 x 10 -18) and strong classification performance (AUC=0.82-0.86), while correlating strongly with protein stability ({Delta}{Delta}G, Spearman=0.9794, P = 5.38 x 10 -28). TSD further reveals PTM sites as topological hubs and neutral variants as evolutionary topological redundancy, enabling a paradigm shift from sequence alignment to geometric dynamics and providing a physics-based biomarker for variants of uncertain significance (VUS). These results upgrade protein mutation modeling from empirical static prediction to provable dynamic mechanism analysis. The source code of this work is publicly available at https://github.com/Harmenlv/LieFold-AI/tree/main.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 17%
10.6%
2
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 3%
9.7%
3
Cell Systems
201 papers in training set
Top 0.4%
7.9%
4
Bioinformatics
1204 papers in training set
Top 3%
7.9%
5
PRX Life
42 papers in training set
Top 0.1%
6.7%
6
PLOS Computational Biology
1863 papers in training set
Top 8%
4.3%
7
Nature Machine Intelligence
70 papers in training set
Top 0.8%
3.5%
50% of probability mass above
8
Advanced Science
286 papers in training set
Top 2%
3.5%
9
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.5%
10
Nature Computational Science
55 papers in training set
Top 0.2%
3.2%
11
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
12
Nucleic Acids Research
1281 papers in training set
Top 6%
2.7%
13
Communications Chemistry
48 papers in training set
Top 0.3%
2.4%
14
Scientific Reports
3612 papers in training set
Top 48%
2.1%
15
Nature Methods
385 papers in training set
Top 4%
2.1%
16
Patterns
78 papers in training set
Top 1%
1.9%
17
Communications Biology
993 papers in training set
Top 12%
1.9%
18
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.1%
19
eLife
5828 papers in training set
Top 57%
1.1%
20
Molecular Systems Biology
162 papers in training set
Top 2%
1.1%
21
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
22
Science Advances
1243 papers in training set
Top 25%
1.1%
23
PNAS Nexus
159 papers in training set
Top 2%
1.0%
24
Biophysical Journal
631 papers in training set
Top 5%
0.8%
25
Cell Reports Methods
165 papers in training set
Top 5%
0.6%
26
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.6%
27
Journal of Molecular Biology
232 papers in training set
Top 4%
0.6%