Back

Surprisal contributes little beyond contextual embeddings in high-gamma ECoG encoding

Sakuma, T.

2026-07-16 neuroscience
10.64898/2026.07.10.737662 bioRxiv
Show abstract

Surprisal and contextual embeddings are both derived from large language models and are widely used to predict neural responses during language comprehension, but it is unclear whether surprisal adds information beyond embeddings. We test this directly: does word surprisal improve out-of-sample prediction of high-gamma ECoG responses after GPT-2 XL contextual embeddings are included? Using public ECoG recordings from natural speech, we fit word-aligned ridge encoding models with baseline stimulus features, GPT-2 XL embeddings, and GPT-2 XL surprisal. Adding one surprisal predictor left held-out correlation essentially unchanged at the center lag, and the effect remained within a prespecified equivalence margin across alternative lags and sensitivity analyses. This near-zero increment suggests that surprisal does not act as an independent predictor. It is better understood as a compressed readout of the same broader predictive state that the embeddings already capture.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.