Back

Multi-source domain generalization with few-shot calibration for cross-dataset EEG state classification under proxy labels

Weng, Z.; Jung, M.

2026-08-24 neuroscience
10.64898/2026.08.19.745846 bioRxiv
Show abstract

Cross-dataset generalization of EEG-based classification under weak, proxy-derived labels remains an open problem for altered-states research. We present a reproducible eight-dataset alignment pipeline that maps eight heterogeneous EEG corpora (712,832 windows; 697,906 with valid labels) to a common 14-channel EPOC+ montage with 63-dimensional spectral features, and we recover the real 1-9 arousal self-assessments for MAHNOB-HCI from session.xml metadata. As a benchmark, Random Forest classifiers are trained on seven source domains and evaluated on the held-out target under both zero-shot and 20%-participant few-shot calibration. The benchmark exposes two concrete methodological pitfalls rather than a performance result: (i) per-class recall shows every target collapsing to a single majority class, and (ii) a within-dataset upper-bound experiment (Table 3) shows that six of eight proxy label sets sit at or below three-class chance even when trained and tested on the same dataset, so the cross-dataset failure is a label-validity problem rather than a transfer-method problem. Across the eight targets (20 seeds, 8,000 evaluation windows per target), zero-shot accuracy averages 36.85% (95% CI 34.40-39.30) and calibrated 43.76% (41.77-45.75), but balanced accuracy stays at 33.01-35.62% (Cohen's kappa <= 0.068), i.e. at chance. The +6.91pp mean change is driven almost entirely by a single target, ds006437 (6.31% -> 60.60%): the median paired change across all 160 seed-pairs is 0.00pp, and after Holm-Bonferroni correction only ds006437 and ds004572 remain significant, the latter with a practically null effect (+0.39pp). Balanced accuracy stays between 33.01% and 35.62% and Cohen's kappa at 0.009 +/- 0.032, i.e. at or barely above three-class chance, while per-class recall shows six of eight targets collapsing to Deep (96.7-100% recall) and two to Light (68.5-99.1%). The collapse persists under SMOTE oversampling, under an EEGNet-v4 deep-learning baseline, and under CORAL and AdaBN feature alignment, which locates the bottleneck in proxy-label validity and class overlap in the feature space rather than in classifier capacity. We position this work as a preliminary methodological study: its contribution is a reproducible eight-dataset alignment pipeline, recovered MAHNOB-HCI arousal self-assessments, a quantitative estimate of split-leakage inflation, and a transparently reported negative result rather than a performance claim.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.