Back

Harmonizing Healthy Cohorts to Support Multicenter Studies on Migraine Classification using Brain MRI Data

Yoon, H.; Schwedt, T. J.; Chong, C. D.; Olatunde, O.; Wu, T.

2023-06-28 health informatics
10.1101/2023.06.26.23291909 medRxiv
Show abstract

Multicenter and multi-scanner imaging studies might be needed to provide sample sizes large enough for developing accurate predictive models. However, multicenter studies, which likely include confounding factors due to subtle differences in research participant characteristics, MRI scanners, and imaging acquisition protocols, might not yield generalizable machine learning models, that is, models developed using one dataset may not be applicable to a different dataset. The generalizability of classification models is key for multi-scanner and multicenter studies, and for providing reproducible results. This study developed a data harmonization strategy to identify healthy controls with similar (homogenous) characteristics from multicenter studies to validate the generalization of machine-learning techniques for classifying individual migraine patients and healthy controls using brain MRI data. The Maximum Mean Discrepancy (MMD) was used to compare the two datasets represented in Geodesic Flow Kernel (GFK) space, capturing the data variabilities for identifying a "healthy core". A set of homogeneous healthy controls can assist in overcoming some of the unwanted heterogeneity and allow for the development of classification models that have high accuracy when applied to new datasets. Extensive experimental results show the utilization of a healthy core. One dataset consists of 120 individuals (66 with migraine and 54 healthy controls) and another dataset consists of 76 (34 with migraine and 42 healthy controls) individuals. A homogeneous dataset derived from a cohort of healthy controls improves the performance of classification models by about 25% accuracy improvements for both episodic and chronic migraineurs. HighlightsO_LIThe harmonization method was established by Healthy Core Construction. C_LIO_LIThe inclusion of a healthy core addresses intrinsic heterogeneity that exists within a healthy control cohort and in multicenter studies. C_LIO_LIThe utilization of a healthy core can increase the accuracy and generalizability of brain imaging-based classification models. C_LIO_LIThe proposed harmonization method offers flexible utilities for multicenter studies. C_LI

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.