A modified decision tree improves generalization across multiple brains proteomic data sets and reveals the role of ferroptosis in Alzheimer disease
Ivanov, M. V.; Kopeykina, A. S.; Kazakova, E. M.; Tarasova, I. A.; Sun, Z.; Postoenko, V. I.; Yang, J.; Gorshkov, M. V.
Show abstract
Low generalization to the patient cohort and variety of experimental conditions in the proteomic search for disease biomarkers are among the main reasons for the bumpy road of quantitative proteomics from discovery stage to clinical validation. Only a small fraction of biomarkers discovered so far by proteomic analysis reaches clinical trials. Here, we presented a machine learning-based workflow for proteomics data analysis, which partially solves some of these issues. In particular, we used a customized decision tree model, which was regulated using a newly introduced parameter, min_cohorts_leaf, that resulted in better generalization of trained models. Further, we analyzed the trend of feature importances curve as a function of min_cohorts_leaf parameter and found that it could be used for accurate feature selection to obtain a list of proteins with significantly improved generalization. Finally, we demonstrated that the recently introduced DirectMS1 search algorithm for protein identification and quantitation provides a simple, yet, a highly efficient solution for the problem of combining multiple data sets obtained using different experimental settings. The developed workflow was tested using five published LC-MS/MS data sets obtained in the large consortia studies of Alzheimers disease brain samples. The selected data sets consist of 535 files in total analyzed using label-free single-shot data-dependent or data-independent acquisitions. Using the proposed modified ExtraTrees model we found that the expressions of two proteins involved in ferroptosis Serotransferrin TRFE and DNA repair nuclease/redox regulator APEX1, are important for explaining a lack of dementia for patients with the presence of neuritic plaques and neurofibrillary tangles.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Isobaric matching between runs and novel PSM-level normalization in MaxQuant strongly improve reporter ion-based quantification 95%
- A Temporal Quantitative Profiling of Newly Synthesized Proteins during Aβ Accumulation 95%
- PELSA-Decipher: a software tool for the processing and interpretation of ligand protein interaction dataset acquired by PELSA 94%
Similar papers in this journal
Similar papers in this journal
- The Cannabis Multi-Omics Draft Map Project 94%
- The hidden secrets of the dental calculus: Calibration of a mass spectrometry protocol for dental calculus protein analysis 92%
- Data independent acquisition mass spectrometry (DIA-MS) analysis of FFPE rectal cancer samples offers in depth proteomics characterization of response to neoadjuvant chemoradiotherapy 92%
Similar papers in this journal
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
- MMS2plot: an R package for visualizing multiple MS/MS spectra for groups of modified and non-modified peptides 94%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 94%
Similar papers in this journal
- NMR Analysis of the Correlation of Metabolic Changes in Blood and Cerebrospinal Fluid in Alzheimer Model Male and Female Mice 93%
- Identification of functionally connected multi-omic biomarkers for Alzheimer’s Disease using modularity-constrained Lasso 93%
- Untargeted saliva metabolomics reveals COVID-19 severity 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.