Back

Calibration-derived decoder discriminability is associated with online P300-speller accuracy, but the fitted mapping does not transport across cohorts

Gorenshtein, A.; Adiniaev, Y.; Liba, T.; Barash, Y.; Klang, E.; Daniel, O.

2026-08-02 neurology
10.64898/2026.07.30.26359353 medRxiv
Show abstract

Objective A calibration-derived score has been related to online P300-speller accuracy in the same session, always within the cohort measured. We evaluated whether a fitted mapping from that score to expected accuracy transports to withheld cohorts. Approach Retrospective secondary analysis of BigP3BCI 1.0.0, the largest published evaluation of this relationship to date: 18 of 20 source studies yielded eligible online outcomes, contributing 271 participants, 739 session-condition records, and 19,611 character selections. Four cohorts document an amyotrophic lateral sclerosis (ALS) population. The predictor was calibration-derived decoder discriminability: cross-validated discriminability of a classifier fitted only to calibration epochs. One source study at a time was withheld from development. Between-cohort variation was summarised by the random-effects standard deviation tau, with participant-clustered standard errors. Main Results The association was positive in all 18 cohorts but varied widely in magnitude and precision (Pearson r 0.190 to 0.928; participant-level r = 0.714, p < 0.001). Pooled estimation error was 0.098 (95% CI 0.091 to 0.107) against a benchmark of 0.146. Calibration did not transport: the intercept had tau 0.87, with a 95% interval for an unrepresented cohort of -1.97 to 1.85 on the log-odds scale, and the slope varied more than tenfold across cohorts (tau 0.43, unrepresented-cohort interval 0.111 to 2.005). A protocol proxy for the stopping rule, median time per selection, reduced the between-cohort slope variance by 66%. The four ALS cohorts alone gave an uncorrected between-cohort slope spread of 0.223, against 0.560 overall. Significance The score carries a reproducible signal about accuracy in the corresponding session, but the mapping between them is cohort-specific: usable for ranking sessions within a setting, not for reporting expected accuracy elsewhere without local recalibration. Few-cohort evaluation, as the ALS subgroup illustrates, understates how much performance varies elsewhere.

Matching journals

The top 12 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.