Back

Cross-Modal Benchmarking of Acoustic Prosody and Ventral Striatal BOLD for Depression-Related Anhedonia Classification: A Pre-Registered Study with the ClinicalWhisper Pipeline

Zhou, C.; Wu, M.; Xiang, Y.; Itti, L.

2026-06-11 neuroscience
10.64898/2026.06.08.728970 bioRxiv
Show abstract

Can computational analysis of a brief voice recording classify depression-related anhedonia as effectively as task-based fMRI? We address this question through a pre-registered cross-modal benchmarking study (osf.io/bsvrj) that evaluates two independent classification pipelines against depression-related anhedonia operationalized via self-report. Anhedonia-- the diminished capacity to experience pleasure or motivation to pursue rewards--is a transdiagnostic marker of reward-system dysfunction that predicts treatment resistance in major depressive disorder, yet current assessment requires either expensive functional neuroimaging or subjective self-report scales, neither of which scales to routine screening. We benchmarked two independent classification pipelines: Stream A extracted 88 acoustic prosody features from DAIC-WOZ clinical interviews (n = 142) using the ClinicalWhisper pipeline (Whisper Large-v3, pyannote diarization, OpenSMILE eGeMAPS v02); Stream B extracted nucleus accumbens BOLD activation during a reward task from the UCLA ds000030 dataset (n = 272). Three classifiers (logistic regression, random forest, gradient-boosted trees) were evaluated under stratified 5-fold cross-validation with fixed, pre-registered hyperparameters. Stream A achieved a best AUC-ROC of 0.63 (random forest; permutation p = .049), confirming H1 that acoustic prosody classifies PHQ-8-defined anhedonia above chance at = .05 (uncorrected), though this result does not survive Bonferroni correction for three classifiers (/3 = .017). Bootstrap analysis of {Delta}AUC confirmed non-inferiority relative to Stream B (H2: 95% CI lower bound > -0.10). However, Stream B itself did not achieve above-chance classification (p = .057), so this non-inferiority finding reflects comparable, modest performance across both modalities rather than equivalence to a validated neural biomarker. Pitch variability features (F0) ranked among the top 5 predictors by SHAP value in the gradient-boosted trees model (H3: partially supported). An exploratory combined model (eGeMAPS + recurrence quantification analysis) reached AUC = 0.65 (GBT; p = .032), though this lift was not statistically significant by DeLong test (z = -0.44, p = .66). These results provide initial evidence that vocal prosody, acquired through a standardized, open-source pipeline, carries depression-related information comparable to fMRI-derived ventral striatal activation for binary classification. However, neither streams operationalization isolates anhedonia from general depressive symptomatology, and the cross-modal design compares different constructs across different cohorts. We refer to the classification target throughout as the "anhedonia composite" to acknowledge that PHQ-8 Items 1+2 conflate anhedonia with dysphoria. We discuss these constraints and their implications for the causal framework motivating this work.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.