Harmonizing Healthy Cohorts to Support Multicenter Studies on Migraine Classification using Brain MRI Data
Yoon, H.; Schwedt, T. J.; Chong, C. D.; Olatunde, O.; Wu, T.
Show abstract
Multicenter and multi-scanner imaging studies might be needed to provide sample sizes large enough for developing accurate predictive models. However, multicenter studies, which likely include confounding factors due to subtle differences in research participant characteristics, MRI scanners, and imaging acquisition protocols, might not yield generalizable machine learning models, that is, models developed using one dataset may not be applicable to a different dataset. The generalizability of classification models is key for multi-scanner and multicenter studies, and for providing reproducible results. This study developed a data harmonization strategy to identify healthy controls with similar (homogenous) characteristics from multicenter studies to validate the generalization of machine-learning techniques for classifying individual migraine patients and healthy controls using brain MRI data. The Maximum Mean Discrepancy (MMD) was used to compare the two datasets represented in Geodesic Flow Kernel (GFK) space, capturing the data variabilities for identifying a "healthy core". A set of homogeneous healthy controls can assist in overcoming some of the unwanted heterogeneity and allow for the development of classification models that have high accuracy when applied to new datasets. Extensive experimental results show the utilization of a healthy core. One dataset consists of 120 individuals (66 with migraine and 54 healthy controls) and another dataset consists of 76 (34 with migraine and 42 healthy controls) individuals. A homogeneous dataset derived from a cohort of healthy controls improves the performance of classification models by about 25% accuracy improvements for both episodic and chronic migraineurs. HighlightsO_LIThe harmonization method was established by Healthy Core Construction. C_LIO_LIThe inclusion of a healthy core addresses intrinsic heterogeneity that exists within a healthy control cohort and in multicenter studies. C_LIO_LIThe utilization of a healthy core can increase the accuracy and generalizability of brain imaging-based classification models. C_LIO_LIThe proposed harmonization method offers flexible utilities for multicenter studies. C_LI
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Brain predictors of fatigue in Rheumatoid Arthritis: a machine learning study 94%
- Application of machine learning and complex network measures to an EEG dataset from ayahuasca experiments 92%
- Computational modelling of the long-term effects of brain stimulation on the local and global structural connectivity of epileptic patients 92%
Similar papers in this journal
Similar papers in this journal
- MRI-based classification of neuropsychiatric systemic lupus erythematosus patients with self-supervised contrastive learning 93%
- Morphometric and Functional Brain Connectivity Differentiates Chess Masters from Amateur Players 92%
- Quantitative evaluation of normal cerebrospinal fluid flow in the Sylvian aqueduct and perivascular spaces of the middle cerebral artery and circle of Willis via 2D phase-contrast MR imaging 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.