Back

A Systematic Evaluation of Machine Learning-based Biomarkers for Major Depressive Disorder across Modalities

Winter, N. R.; Blanke, J.; Leenings, R.; Ernsting, J.; Fisch, L.; Sarink, K.; Barkhau, C.; Thiel, K.; Flinkenflügel, K.; Winter, A.; Goltermann, J.; Meinert, S.; Dohm, K.; Repple, J.; Gruber, M.; Leehr, E. J.; Opel, N.; Grotegerd, D.; Redlich, R.; Nitsch, R.; Bauer, J.; Heindel, W.; Gross, J.; Andlauer, T. F. M.; Forstner, A. J.; Nöthen, M. M.; Rietschel, M.; Hofmann, S. G.; Pfarr, J.-K.; Teutenberg, L.; Usemann, P.; Thomas-Odenthal, F.; Wroblewski, A.; Brosch, K.; Stein, F.; Jansen, A.; Jamalabadi, H.; Alexander, N.; Straube, B.; Nenadic, I.; Kircher, T.; Dannlowski, U.; Hahn, T.

2023-02-27 psychiatry and clinical psychology
10.1101/2023.02.27.23286311 medRxiv
Show abstract

BackgroundBiological psychiatry aims to understand mental disorders in terms of altered neurobiological pathways. However, for one of the most prevalent and disabling mental disorders, Major Depressive Disorder (MDD), patients only marginally differ from healthy individuals on the group-level. Whether Precision Psychiatry can solve this discrepancy and provide specific, reliable biomarkers remains unclear as current Machine Learning (ML) studies suffer from shortcomings pertaining to methods and data, which lead to substantial over-as well as underestimation of true model accuracy. MethodsAddressing these issues, we quantify classification accuracy on a single-subject level in N=1,801 patients with MDD and healthy controls employing an extensive multivariate approach across a comprehensive range of neuroimaging modalities in a well-curated cohort, including structural and functional Magnetic Resonance Imaging, Diffusion Tensor Imaging as well as a polygenic risk score for depression. FindingsTraining and testing a total of 2.4 million ML models, we find accuracies for diagnostic classification between 48.1% and 62.0%. Multimodal data integration of all neuroimaging modalities does not improve model performance. Similarly, training ML models on individuals stratified based on age, sex, or remission status does not lead to better classification. Even under simulated conditions of perfect reliability, performance does not substantially improve. Importantly, model error analysis identifies symptom severity as one potential target for MDD subgroup identification. InterpretationAlthough multivariate neuroimaging markers increase predictive power compared to univariate analyses, single-subject classification - even under conditions of extensive, best-practice Machine Learning optimization in a large, harmonized sample of patients diagnosed using state-of-the-art clinical assessments - does not reach clinically relevant performance. Based on this evidence, we sketch a course of action for Precision Psychiatry and future MDD biomarker research. FundingThe German Research Foundation, and the Interdisciplinary Centre for Clinical Research of the University of Munster.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.