Back

Sequence determinants of pathogenicity in glucose-6-phosphatase linked to glycogen storage disease type 1a

Stein, R. A.; Hawes, E. M.; Norphlet, C. M.; Rakonick, M. H.; Harris, S. A.; Sivam, T.; Lucerne, A. M.; Da Silva, V. R.; O'Brien, R. M.; Claxton, D. P.

2026-07-28 biochemistry
10.64898/2026.07.27.741017 bioRxiv
Show abstract

Glycogen storage disease type 1a (GSD1a) is an autosomal recessive Mendelian disorder that can be caused by missense variants in glucose-6-phosphatase catalytic subunit 1 (G6PC1). Although hundreds of missense variants have been identified, the vast majority are of unknown clinical significance, and the molecular mechanism(s) of bona fide pathogenic variants are ill-defined. We combine bioinformatic data and clinical associations with the protein language model AlphaMissense to guide mechanistic exploration of 78 missense variants at 55 residue positions using robust biochemical and biophysical assays to distill general principles of enzyme dysfunction. Correlation analysis established a strong linear relationship between folded G6PC1 abundance and catalytic capacity for most variants. Pathogenic variants within this paradigm were linked to compromised stability and activation of the unfolded protein response. However, outliers characterized by relatively high abundance, yet low activity clustered to a network of sidechains adjacent to the active site that allosterically modulate catalysis. Contextualized by recent high-resolution structures and AlphaFold modeling, our holistic analysis of G6PC1 in vitro metrics facilitates clinical (re)classification of variants according to explicit molecular phenotypes and identifies therapeutic design directions. Moreover, our approach illustrates a blueprint for variant characterization that integrates computational prediction with experimental validation to discover disease etiology.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.