Proteins as Statistical Languages: Information-Theoretic Signatures of Proteomes Across the Tree of Life
Alegre, E. O. T.
Show abstract
Protein sequences are commonly interpreted through biochemical and evolutionary lenses, emphasizing structure-function relationships and selection in sequence space. Here we develop a complementary viewpoint: proteins as statistical languages--strings over a finite alphabet generated by constrained stochastic processes. We formalize intrinsic informational descriptors of protein ensembles, including composition entropy H1, adjacent mutual information I1, and separation-dependent information profiles Id. A null-model ladder (uniform, composition-matched i.i.d., and Markov-1) separates compositional effects from genuine positional dependence. We then evaluate these descriptors empirically across 20 UniProt reference proteomes spanning major clades, using protein-level bootstrap resampling and matched synthetic controls. Real proteomes consistently depart from composition-matched i.i.d. baselines and exhibit information profiles that remain elevated beyond the decay expected under first-order Markov surrogates, indicating dependencies beyond local transition statistics. Finally, a compressibility proxy (gzip) provides an orthogonal signature of redundancy relative to i.i.d. controls at matched composition. Together, these results support the view of proteomes as constrained statistical languages and provide model-agnostic fingerprints for comparing sequence ensembles.These signatures provide a lightweight diagnostic layer for comparing proteomes prior to mechanistic modeling
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Group-walk, a rigorous approach to group-wise false discovery rate analysis by target-decoy competition 94%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 93%
- The Ortholog Conjecture Revisited: the Value of Orthologs and Paralogs in Function Prediction 93%
Similar papers in this journal
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 92%
- Global transcription regulation revealed from dynamical correlations in time-resolved single-cell RNA-sequencing 92%
- Studying stochastic systems biology of the cell with single-cell genomics data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.