Multi-source domain generalization with few-shot calibration for cross-dataset EEG state classification under proxy labels
Weng, Z.; Jung, M.
Show abstract
Cross-dataset generalization of EEG-based classification under weak, proxy-derived labels remains an open problem for altered-states research. We present a reproducible eight-dataset alignment pipeline that maps eight heterogeneous EEG corpora (712,832 windows; 697,906 with valid labels) to a common 14-channel EPOC+ montage with 63-dimensional spectral features, and we recover the real 1-9 arousal self-assessments for MAHNOB-HCI from session.xml metadata. As a benchmark, Random Forest classifiers are trained on seven source domains and evaluated on the held-out target under both zero-shot and 20%-participant few-shot calibration. The benchmark exposes two concrete methodological pitfalls rather than a performance result: (i) per-class recall shows every target collapsing to a single majority class, and (ii) a within-dataset upper-bound experiment (Table 3) shows that six of eight proxy label sets sit at or below three-class chance even when trained and tested on the same dataset, so the cross-dataset failure is a label-validity problem rather than a transfer-method problem. Across the eight targets (20 seeds, 8,000 evaluation windows per target), zero-shot accuracy averages 36.85% (95% CI 34.40-39.30) and calibrated 43.76% (41.77-45.75), but balanced accuracy stays at 33.01-35.62% (Cohen's kappa <= 0.068), i.e. at chance. The +6.91pp mean change is driven almost entirely by a single target, ds006437 (6.31% -> 60.60%): the median paired change across all 160 seed-pairs is 0.00pp, and after Holm-Bonferroni correction only ds006437 and ds004572 remain significant, the latter with a practically null effect (+0.39pp). Balanced accuracy stays between 33.01% and 35.62% and Cohen's kappa at 0.009 +/- 0.032, i.e. at or barely above three-class chance, while per-class recall shows six of eight targets collapsing to Deep (96.7-100% recall) and two to Light (68.5-99.1%). The collapse persists under SMOTE oversampling, under an EEGNet-v4 deep-learning baseline, and under CORAL and AdaBN feature alignment, which locates the bottleneck in proxy-label validity and class overlap in the feature space rather than in classifier capacity. We position this work as a preliminary methodological study: its contribution is a reproducible eight-dataset alignment pipeline, recovered MAHNOB-HCI arousal self-assessments, a quantitative estimate of split-leakage inflation, and a transparently reported negative result rather than a performance claim.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Do try this at home: Age prediction from sleep and meditation with large-scale low-cost mobile EEG 92%
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 92%
- Harmonizing and aligning M/EEG datasets with covariance-based techniques to enhance predictive regression modeling 91%
Similar papers in this journal
- Automated EEG mega-analysis II: Cognitive aspects of event related features 93%
- Class imbalance should not throw you off balance: Choosing the right classifiers and performance metrics for brain decoding with imbalanced data 93%
- A reusable benchmark of brain-age prediction from M/EEG resting-state signals 92%
Similar papers in this journal
- An infection prediction model developed from inpatient data can predict out-of-hospital COVID-19 infections from wearable data when controlled for dataset shift 92%
- Recurrent Neural Network-based Acute Concussion Classifier using Raw Resting State EEG Data 92%
- Alpha activity neuromodulation induced by individual alpha-based neurofeedback learning in ecological context: A double-blind randomized study 91%
Similar papers in this journal
- Family Lexicon: using language models to encode memories of personally familiar and famous people and places in the brain 92%
- Towards a real-world brain-computer interface for image retrieval 92%
- Discrimination of sleep and wake periods from a hip-worn raw acceleration sensor using recurrent neural networks 92%
Similar papers in this journal
- Decoupling simultaneous motor imagination and execution via orthogonal ECoG neural representations 90%
- Spatiotemporal brain hierarchies of auditory memory recognition and predictive coding 89%
- Transcranial Focused Ultrasound to V5 Enhances Human Visual Motion Brain-Computer Interface by Modulating Feature-Based Attention 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.