Exploring Machine Learning Models to Uncover Pathways in ALS Pathogenesis Using Immunohistochemical Features
Kuruvilla, J. M.; Rifai, O. M.; Longden, J.; Gregory, J. M.; Vallejo, M.
Show abstract
Amyotrophic Lateral Sclerosis (ALS) is a degenerative disease of motor neurons that leads to muscle wasting, paralysis, and death, with an average life expectancy of 2-5 years. Approximately 10-15% of ALS cases are familial (fALS), typically linked to, but not always caused by identifiable inherited genetic mutations. The remaining 85-90% are considered sporadic ALS (sALS), which typically occurs without a clear family history. It is thought to result from a combination of genetic and non-genetic risk factors. ALS imposes heavy physical, psychological, and financial burdens on patients and caregivers. Early diagnosis is critical but remains challenging due to clinical variability and overlapping symptoms with other motor neuron disorders. Current diagnostic methods, including genetic testing and neurophysiological techniques, face limitations in reproducibility and accessibility, while machine learning offers potential by detecting patterns that traditional methods overlook. This study applies machine learning to characterise disease status in C9orf72-ALS patients and evaluate how pathological biomarkers relate to disease mechanisms. A tabular dataset from post-mortem brain tissue of 10 C9orf72-ALS patients and 10 controls was used to train models and benchmark results against Rifai et al. (2022). Models included random forest, support vector machine, xgboost, logistic regression, artificial neural networks, and ensembles, validated using 3-fold and 5-group cross-validation. The best model result was of random forest with 3-fold cross-validation for Iba1, achieving 88% sensitivity (p = 0.0011) and 83% specificity (p = 0.0004). However, as 3-fold cross-validation is less robust, we expect more reliable and stable results from 5-fold grouped cross-validation and, in future, repeated cross-validation approaches. Machine learning offers insights into ALS, with implications for potential patient stratification and case identification.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 95%
- Association of Graph-based Spatial Features with Overall Survival Status of Glioblastoma Patients 93%
- Identifying relationships between imaging phenotypes and lung cancer-related mutation status: EGFR and KRAS 93%
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 93%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 93%
- Predicting the physiological effects of multiple drugs using electronic health record 92%
Similar papers in this journal
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 91%
- Comparing randomized trial designs to estimate treatment effect in rare diseases with longitudinal models: a simulation study showcased by Autosomal Recessive Cerebellar Ataxias using the SARA score 91%
- Prediction-powered Inference for Clinical Trials 89%
Similar papers in this journal
- An Inexpensive Smartphone-Based Device and Predictive Models for Rapid, Non-Invasive, and Point-of-Care Monitoring of Ocular and Cardiovascular Complications Related to Diabetes 91%
- Generalizable electroencephalographic classification of Parkinson’s Disease using deep learning 91%
- End-to-End Machine Learning based Discrimination of Neoplastic and Non-neoplastic Intracerebral Hemorrhage on Computed Tomography 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.