Back

When Meaning Matters Most: Rethinking Cloze Probability in N400 Research

Arkhipova, Y.; Lopopolo, A.; Vasishth, S.; Rabovsky, M.

2025-05-02 neuroscience
10.1101/2025.04.29.651301 bioRxiv
Show abstract

The N400 component of ERPs is modulated by how predictable a word is, but predictability is usually quantified with lexical cloze--the probability that readers supply that exact word in offline sentence completion tasks. This form-based metric is at odds with decades of evidence that the N400 is primarily sensitive to meaning. Here, we asked whether a measure of semantic feature predictability can better account for N400 amplitude modulation. We reanalysed two independent EEG datasets (N = 26 and N = 334), computing lexical and semantic cloze for each critical word. Across both datasets, semantic cloze emerged as a better predictor of the N400 data than lexical cloze. Using the same materials, we then compared semantic and lexical cloze with probabilities from four large language models (GPT-2, GPT-2.7b, RoBERTa, ALBERT). None of the LLM-derived predictors outperformed semantic cloze. Our findings support the view that the N400 primarily reflects semantic--not exact-word--processing. Methodologically, we argue that replacing lexical cloze with semantic cloze can substantially increase the explanatory power of N400 studies, and caution against substituting human norms with raw LLM probabilities.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.