Protein Compositional Ratio Representation (PCRR)Systematically Improves Human Disease Prediction
Madduri, A. V.; Ellis, R. J.; Patel, C. J.
Show abstract
Plasma proteomics captures a functional snapshot of human physiology; yet, most machine learning models treat protein abundances as independent variables, ignoring the fact that biological systems and proteomic measurements are inherently compositional. Many molecular processes depend not on absolute concentrations but on relative balances: receptor-ligand stoichiometry, enzyme-substrate ratios, and homeostatic feedbacks that govern signaling and metabolism. We propose that these relationships are best captured through pairwise protein ratios, which more faithfully reflect underlying biochemical constraints than raw expression values. We evaluate a machine learning framework that models pairwise log-ratios of proteins (log(A) -log(B)) as features, thereby encoding compositional structure directly into the learning space. Applied to the ROSMAP plasma proteomics cohort (n = 871), this approach substantially improved the classification of Alzheimers subtypes (NCI, MCI, AD, AD+) with an average AUROC gain of +0.1274 over a strong baseline that incorporated raw proteomics and demographics. The top-ranked ratios (e.g., SEMA3C:TMEM70, IDUA:NPTXR) captured converging pathogenic pillars of Alzheimers disease, including microglial activation, proteostasis dysregulation, and lipid-clearance imbalance, highlighting that ratio-based features recover biologically coherent axes of disease. To assess generality, we scaled the method to the UK Biobank proteomic dataset (n > 53,000; 587 phenotypes). The ratio-based model outperformed raw-level models in 95.1 % of diseases, with statistically significant (FDR < 0.05) gains in 56.7 %. Together, these results suggest that proteomic data should be viewed and modeled as compositional systems, where relative protein abundances carry the accurate functional signal. This insight supports the broader utility of ratio-based representations for disease prediction and biomarker discovery.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SomaModules: a pathway enrichment approach tailored to SomaScan data 93%
- Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins 93%
- Bridging Simplicity and Depth in Single-Cell Proteomics: A Cost-Effective Workflow and Expanded Framework for Data Evaluation 93%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 94%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 92%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 92%
Similar papers in this journal
- Cross-ancestry information transfer framework improves protein abundance prediction and protein-trait association identification 92%
- eNODAL: an experimentally guided nutriomics data clustering method to unravel complex drug-diet interactions 92%
- Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.