Novel Pipelines to Extract Differences in Proteome Dynamics Based on Health Status
Xu, B.; Zhao, J.; Huang, T.; Honfo, S. H.; Trumpff, C.; Picard, M.; Cohen, A.; Liu, M.
Show abstract
Understanding dynamics and co-regulatory patterns in the human proteome is a promising path for unraveling the molecular basis of health and disease. Nevertheless, there remains an open challenge in extracting concise information from high-throughput proteomic data that can effectively characterize and predict health. We develop novel statistical and computational pipelines to tackle this problem in a longitudinal saliva proteomics data set collected throughout the awakening response in six healthy controls and six subjects with severe mitochondrial disease (MitoD), a clinical condition caused by genetic mitochondrial defects that affects cellular energy transformation and alters multiple dimensions of health. We undertook three independent unsupervised approaches to characterize proteome dynamics and assessed their ability to separate MitoD individuals from controls. First, we designed a permutation test to detect the global difference in the proteomic co-regulation structure between healthy and unhealthy subjects. Second, we performed non-linear embedding and cluster analysis on elasticity to capture a more complicated relationship between health and the proteome. Third, we developed a machine learning algorithm to extract low-dimensional representations of the proteome dynamic and use them to cluster subjects into healthy and unhealthy groups without any knowledge of their true status. All three methods showed clear differences between MitoD individuals and controls. Our results revealed a significant and consistent association between MitoD status and the saliva proteome at multiple levels during the awakening response, including its dynamic change, co-regulation structure, and elasticity. This connection is not restricted to a few MitoD-specific proteins but spreads over a wide range of proteins from many body functions and pathways. Pipelines such as those shown here are the first step toward establishing interpretable and accurate prediction rules for health based on proteome dynamics.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SomaModules: a pathway enrichment approach tailored to SomaScan data 94%
- Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins 93%
- Effects of in vitro hemolysis and repeated freeze-thaw cycles in protein abundance quantification using the SomaScan and Olink assays 92%
Similar papers in this journal
Similar papers in this journal
- eNODAL: an experimentally guided nutriomics data clustering method to unravel complex drug-diet interactions 95%
- Deep-learning enables proteome-scale identification of phase-separated protein candidates from immunofluorescence images 91%
- Single-cell multi-omics and spatial multi-omics data integration via dual-path graph attention auto-encoder 91%
Similar papers in this journal
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 94%
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 93%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.