Leakage-Aware Decoding of Music Perception and Cued Imagery Across the Full OpenMIIR EEG Cohort: A Reproducible Analysis of the Generalization Boundary
Wang, Y.; Wang, K.
Show abstract
Music perception and musical imagery provide a controlled setting for studying whether scalp electroencephalography (EEG) captures reproducible differences between externally driven and internally generated auditory states. We tested whether the public OpenMIIR dataset supports leakage-aware decoding of music perception versus cued musical imagery across its full ten-subject cohort, and we characterized the boundary beyond which the decoded signal fails to generalize. Using compact spectral and temporal EEG features, we applied stratified trial-grouped cross-validation, dummy and shuffled-label negative controls, leave-one-subject-out (LOSO) testing, a 1000-fold trial-level label-permutation test, and a group-level one-sided Wilcoxon test over per-subject within-subject accuracies, with Benjamini-Hochberg (BH) correction across the family of tested hypotheses. Within subjects, decoding was above chance at the population level: a group Wilcoxon test on logistic-regression accuracy gave p = 0.0195 with a large effect size (Cohens dz = 0.95; 7 of 10 subjects above chance), confirmed by a pooled trial-level permutation test (p = 0.0040). Pooled trial-grouped balanced accuracy reached 0.567 [0.550,0.586] for random forest and 0.543 [0.523,0.561] for logistic regression, exceeding both dummy and shuffled-label controls. Cross-subject transfer was weaker and model-dependent: under LOSO, random forest reached 0.559 [0.523, 0.597], above its dummy baseline (uncorrected p = 0.014), whereas logistic regression did not generalize (0.518, p = 0.165). Under BH correction across the nine tested hypotheses, six comparisons survived at q < 0.05 (smallest q = 0.036), all involving the permutation test or the nonlinear model, while the linear models cross-subject contrasts did not. These results indicate that OpenMIIR EEG supports modest but reproducible within-subject discrimination of music perception and cued imagery, with a linear-within-subject versus nonlinear-cross-subject generalization boundary, and they show why public EEG music data require leakage-aware validation and calibrated subject-generalization claims.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A supervised data-driven spatial filter denoising method for speech artifacts in intracranial electrophysiological recordings 93%
- Probing for Intentions: The Early Readiness Potential does not Reflect Awareness of Motor Preparation 93%
- All spectral frequencies of neural activity reveal semantic representation in the human anterior ventral temporal cortex 92%
Similar papers in this journal
- The impact of musical expertise on disentangled and contextual neural encoding of music revealed by generative music models 94%
- Spatiotemporal brain hierarchies of auditory memory recognition and predictive coding 94%
- Timing along the cardiac cycle modulates neural signals of reward-based learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.