Predicting Huntington's disease state with ensemble learning & sMRI: more than just the striatum
Kohli, M.; Pustina, D.; Warner, J. H.; Alexander, D. C.; Scahill, R. I.; Tabrizi, S. J.; Sampaio, C.; Wijeratne, P. A.
Show abstract
Developing effective treatments for Huntingtons disease (HD) requires reliable markers of disease progression. Striatal atrophy has been the hallmark of HD progression, but volumetric anomalies are also found in other brain regions. Little is known about the potential increase in predictive biomarking accuracy when volumetric scores from multiple brain regions are combined to predict the HD status of individual participants. We used cross-sectional structural MRI data from 184 HD gene-positive participants to a) test a novel ensemble machine learning model in classifying participants in one of four HD progression states (PreHD A; PreHD B; HD1; HD2), and (b) identify the brain regions that carry HD biomarking signal from 15 regions. We used 5-fold cross validation and backward feature elimination to find the optimal predictors and investigated the stability of the findings through repeated analyses. The ensemble predictive model systematically matched or outperformed the accuracy of nine standard machine learning models, reaching 55.3%{+/-}6.1 balanced accuracy in 4-group classification. The accuracy was higher for binary classifications (PreHD vs HD: 83.3%{+/-}6.3; PreHD A vs PreHD B: 76.7%{+/-}8.0; PreHD B vs HD1: 75.9%{+/-}8.5; HD1 vs HD2: 70.9%{+/-}9.4). Striatal structures (caudate and putamen) were systematically found to be top predictors. However, the accuracy increased substantially when we included other regions in the model (e.g., occipital cortex, lateral ventricles, cingulate, temporal lobe). Optimal models frequently included 2-7 brain regions from different areas. Overall, the accuracy of classifications remained stable across repetitions but the list of selected brain regions could vary, likely due to collinearities in volumetric scores. This is the first study to demonstrate the improvement of classification accuracy when predicting HD progression with a stacked ensemble model. Our findings indicate that HD progression is marked not only by striatal atrophy but also by volumetric changes outside the striatum, without which biomarking models cannot achieve optimal results. The robust methods applied here exposed instability in the selection of brain regions despite the sizeable sample size (n=184); such instabilities could lead to different conclusions in different studies when single analyses are applied on smaller sample sizes. From a translational perspective, our study informs on the selection of candidate endpoints or target regions for therapeutic intervention in future clinical trials.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Apathy progression is associated with brain atrophy and white matter damage in Parkinson's disease 92%
- The sequence of regional structural disconnectivity due to multiple sclerosis lesions 92%
- Machine learning-based prediction of motor status in glioma patients using diffusion MRI metrics along the corticospinal tract 92%
Similar papers in this journal
- Characterizing resting-state EEG oscillatory and aperiodic activity in neurodegenerative diseases: A multicentric study 91%
- Segmentation of the Human Tongue Musculature Using MRI: Field Guide and Validation in Motor Neuron Disease 91%
- Alzheimer Disease Knowledge Graph Enhances Knowledge Discovery and Disease Prediction 89%
Similar papers in this journal
Similar papers in this journal
- Multi-compartment analysis of the complex gradient-echo signal quantifies myelin breakdown in premanifest Huntington's disease 93%
- Large Scale Functional and Effective Connectivity Alterations cross the Huntington's Disease Integrated Staging System 93%
- Deformation Based Morphometry Study Of Longitudinal MRI Changes In Behavioral Variant Frontotemporal Dementia 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.