A Novel Algorithm for the Harmonization of Pan-cancer Proteomics
Simchi, N.; Givton, O.; Rinberg, J.; Shtrikman, A.; Geiger, T.; Nesvizhskii, A. I.; Seger, E.; Pevzner, K.
Show abstract
Proteomic characterization of cancer tissues holds the potential to advance therapeutic options and reveal novel biomarkers by unlocking insights available only on the proteome level. However, proteomics data analysis is greatly challenged by systematic technical variability in experimental protocols, instrumentation and data processing, restricting comparisons between studies. With the continued and unprecedented growth of proteomics datasets, a comprehensive strategy for harmonizing these datasets is necessary to enable large-scale integrative analyses. Herein, we describe a novel framework for pan-cancer harmonization and imputation, which offers the scientific community an updated approach to this challenge. Rather than relying on a single batch-effect correction algorithm, our multi-step approach accurately addresses critical systematic differences with custom-tailored solutions, including standardized reanalysis of raw data and an autoencoder for pan-cancer integration. By introducing a suite of benchmarks, we bridge the critical gap in reliable harmonization evaluation. Using this framework, we created a harmonized pan-cancer dataset and demonstrated its superiority over existing solutions and previous pan-cancer harmonization efforts. We further demonstrated its utility in revealing prognostic markers, estimating indication-wide biomarker prevalence, and facilitating target discovery for cancer subtypes. We expect our work to provide powerful tools supporting proteomics research for precision cancer medicine.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 96%
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 95%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 95%
Similar papers in this journal
Similar papers in this journal
- Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins 95%
- Analysis and visualization of quantitative proteomics data using FragPipe-Analyst 95%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 95%
Similar papers in this journal
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
- STEAM: Spatial Transcriptomics Evaluation Algorithm and Metric for clustering performance 94%
- Cross-ancestry information transfer framework improves protein abundance prediction and protein-trait association identification 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.