Back

Discovering Biomarker Proteins and Peptides for Parkinson's Disease Prognosis Prediction with Machine Learning and Interpretability Methods

Park, H.-m.; Kabanga, E.; Moon, D.-i.; Chung, M.; Im, J.; Kim, Y.; Van Messem, A.; De Neve, W.

2023-05-22 bioinformatics
10.1101/2023.05.18.541380 bioRxiv
Show abstract

Parkinsons disease is a neurodegenerative disorder that affects millions of people worldwide, posing significant challenges for diagnosis and treatment. This study presents a machine learning pipeline for identifying candidate biomarker proteins and peptides from cerebrospinal fluid mass spectrometry (CSF-MS) tests in Parkinsons disease patients. Our pipeline comprises two main stages: (1) model training using mutual information-based feature selection and five different machine learning regressors and (2) identification of candidate biomarkers by combining three types of interpretability methods. Our regression models demonstrated promising effectiveness in predicting the Movement Disorder Society-Unified Parkinsons Disease Rating Scale (MDS-UPDRS) scores, with UPDRS-1 receiving the best predictions, followed by UPDRS-3 and UPDRS-2. Furthermore, our pipeline identified 11 proteins and peptides as potential biomarkers for Parkinsons disease, excluding Levodopa usage which trivially has the most significant impact on the prognosis prediction. Comparisons with four additional pipelines confirmed the effectiveness of our approach in terms of both model performance and biomarker identification. In conclusion, our study presents a comprehensive machine learning pipeline that demonstrates effectiveness in predicting the severity of Parkinsons disease using CSF-MS tests. Our approach also identifies potential biomarkers, which could aid in the development of new diagnostic tools and treatments for patients with Parkinsons disease.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.