BiomarkerML: A cloud-based proteomics ML workflow for biomarker discovery
Zhou, Y.; Maurya, A. K.; Deng, Y.; Fletcher, M. P.; Ren, C.; Taylor, A.
Show abstract
BackgroundHigh-throughput affinity and mass-spectrometry-based proteomic studies of large clinical cohorts generate high-dimensional proteomic data useful for accelerated disease biomarker discovery. A powerful approach to realizing the potential of these big, complex, and non-linear data, whilst ensuring reproducible results, is to use automated machine learning (ML) and deep learning (DL) pipelines for their analysis. However, there remains a gap in comprehensive ML workflows tailored to proteomic biomarker discovery and designed for biomedical researchers who need pipelines to optimally self-configure and automatically avoid over-fitting. FindingsWe present BiomarkerML, a cloud-based workflow for automated, reproducible, and efficient ML/DL analysis of proteomic data for biomarker discovery, designed for novice-ML users and implemented in Python, R and Workflow Description Language (WDL). BiomarkerML: ingests proteomic and clinical data alongside sample labels; pre-processes data for model fitting and optionally performs dimensionality reduction and visualization; fits a catalogue of ML and DL classification and regression models; and calculates model performance metrics for model comparison. Next, the workflow applies mean SHapley Additive exPlanations (SHAP) to quantify the contribution of each protein to model predictions across all samples. Finally, proteins with high mean SHAP values, and their co-expressed protein network interactors, are identified as candidate biomarkers. Importantly, hyperparameters - configuration variables set prior to training models - are automatically fine-tuned via grid-search, and BiomarkerML employs weighted, nested cross-validation to avoid model over-fitting and data leakage. ConclusionsBiomarkerML is scalable, provides a standardized, user-friendly interface, and streamlines analyses to ensure reproducibility of results. Overall, BiomarkerML is a significant advancement, enabling novice-ML researchers to use cutting-edge ML/DL tools to identify disease biomarkers in complex proteomic data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 95%
- Tail-Robust Quantile Normalization 94%
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 94%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 95%
- IMMUNOTAR - Integrative prioritization of cell surface targets for cancer immunotherapy 94%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 94%
Similar papers in this journal
- nf-encyclopedia: A cloud-ready pipeline for chromatogram library data-independent acquisition proteomics workflows 95%
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 95%
- Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.