Machine learning cross-platform proteomic imputation enables protein quality scoring and replication of epidemiological associations
Li, L.; Alaa, A.; Tan, Y.; Demirel, I.; Friedman, S.; Zha, Q.; Trac, R. P.; Taylor, K. D.; Yu, B.; Ballantyne, C. M.; Deo, R.; Dubin, R.; Tsai, M. Y.; Peloso, G. M.; Brody, J.; Austin, T.; Psaty, B. M.; Nicholas, J.; Raffield, L. M.; Tahir, U.; Coresh, J.; Hornsby, W.; Chan, A.; Rich, S. S.; Rotter, J. I.; Ganz, P.; Gerszten, R.; Philippakis, A.; Natarajan, P.; Yu, Z.
Show abstract
High-throughput affinity-based proteomics has advanced biomedical research, yet fundamental, persistent discordance between mainstream platforms (SomaScan and Olink) routinely undermines the replication of findings. This platform-driven non-replication complicates downstream biological validation and biomarker prioritization. Here, we develop a machine learning-based framework for cross-platform protein value imputation to resolve this translational bottleneck. Using paired proteomic data measured by both SomaScan and Olink from 5,325 participants of the Multi-Ethnic Study of Atherosclerosis, we developed models to impute cross-platform measurements and applied them to two independent and demographically distinct cohorts (Cardiovascular Health Study [N=3,171] and UK Biobank [UKB; N=41,405]) for external validation. Our bi-directional model 1) established an imputation performance-based protein fidelity index, validated against gold-standard measurements from Atherosclerosis Risk in Communities study (N=101) and Nurses Health Study (N=54), 2) enabled imputation of platform-exclusive protein measurements, and 3) facilitated calibration of overlapping proteins. We demonstrate the utility of this framework through three applications: 1) fidelity-informed analyses enhanced the replication of biomarker discovery, 2) recovery of SomaScan signals that were previously inaccessible in UKBs original Olink measurements, and 3) improved replication performance for overlapping proteins. Our study offers a translational roadmap that allows researchers to achieve reliable epidemiological replication, target specific assays for future optimization, and prioritize biological signal over platform noise.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Turnover and replication analysis by isotope labeling (TRAIL) reveals the influence of tissue context on protein and organelle lifetimes 94%
- Causal integration of multi-omics data with prior knowledge to generate mechanistic hypotheses 94%
- Machine learning-guided deconvolution of plasma protein levels 94%
Similar papers in this journal
- Cross-ancestry information transfer framework improves protein abundance prediction and protein-trait association identification 96%
- MetalinksDB: a flexible and contextualizable resource of metabolite-protein interactions 93%
- Deep Learning-based Pseudo-Mass Spectrometry Imaging Analysis for Precision Medicine 92%
Similar papers in this journal
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 94%
- scFeatures: Multi-view representations of single-cell and spatial data for disease outcome prediction 94%
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.