Neurobiology of Language
● MIT Press
All preprints, ranked by how well they match Neurobiology of Language's content profile, based on 29 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Klein, C. C.; Berger, P.; Wiesmann, C. G.; Friederici, A. D.
Show abstract
In preschool years, children take important steps in grammar acquisition, which are essential to learning their native language. A central aspect is the acquisition of the morpho-syntactic rule system, which forms an intersection between words and sentences. In adults, rule-based linguistic processes are supported by the dorsal fiber pathway to BA44, the arcuate fascicle. This pathway matures relatively late in development, raising the question of whether it already supports grammar processes in the early preschool years, or whether early grammar acquisition is supported by different, earlier-maturing fiber pathways. In two independent samples of 3- to 5-year-old children (N = 90 and N = 30), we examined the association between the maturation of fiber pathways of the language network and childrens noun plural assignment as an index of their morpho-syntactic abilities. This revealed consistent differences between 3-year-olds and 4- to 5-year-olds. The 4- and 5-year-olds, but not 3-year-olds, showed a relation of morpho-syntax with both the dorsal pathway to BA44, supporting syntactic processes, and the dorsal pathway to BA6, supporting phonological processes in adults. Our results suggest that, in contrast to adults, preschool-aged children rely on both dorsal fiber pathways for morpho-syntax. This difference might point to different processing strategies reflecting the transition from phonology-based statistical learning to rule-based learning in grammar acquisition.
Wang, C.; Fan, Z.; Han, Z.; Bi, Y.; Li, J.
Show abstract
Recent large language models (LLMs) have demonstrated remarkable proficiency in complex linguistic tasks and have been shown to share certain computational principles with human language processing. However, whether LLMs internal components perform distinct functions, like semantic and syntactic processing in human language systems, remains unclear. Here, we systematically disrupted components of LLMs to simulate the behavioral profiles of aphasia--a disorder characterized by specific language deficits resulting from brain injury. Our findings showed that lesioning specific components of LLMs could replicate behaviors characteristic of different aphasia subtypes. Notably, while semantic deficits as those observed in Wernickes and Conduction aphasia, were relatively straightforward to simulate, reproducing syntactic and lexical impairments, as seen in Brocas and Anomic aphasia, proved more challenging. Together, these results highlight both parallels and discrepancies between emergent modularity in LLMs and the human language system, providing new insights into how information is represented and processed in artificial and biological intelligence.
Dunagan, D.; Low, D. S.; Yue, S.; Meyer, L.; Hale, J.
Show abstract
Human sentence comprehension proceeds word-by-word, with prior research proposing two central sources of cognitive demand during incremental processing: forward-looking disambiguation of the incoming information stream, and backward-looking retrieval of information associated with previous words from working memory. Recent work has shown that Transformer-based language models successfully generate predictions about sentence processing load in human psycho- and neurolinguistic data by operationalizing disambiguation cost as next-token surprisal, and memory retrieval cost as normalized attention entropy (NAE). Such models, however, remain difficult to interpret as it is not well understood what factors play causally into the decision to assign a cost value to a given word in such artificial neural networks. Here, we present interpretable and cognitively grounded models of disambiguation and memory retrieval and evaluate their neural alignment and spatio-temporal correlates using human magnetoencephalography responses to naturalistic narrative speech. Multivariate temporal response function modeling demonstrates firstly that these human-bias-informed models fare equally well in accounting for observed human language processing data as their Transformer counterparts. This same modeling framework then suggests that surprisal and NAE temporally dissociate in the cortical language network -- surprisal being predictive of bilateral superior temporal gyrus and supramarginal gyrus activation [~]300-500 ms, and NAE being predictive of activity in the same regions, but later [~]750-850 ms. By demonstrating that interpretable neurocomputational models can achieve meaningful brain alignment while maintaining explanatory transparency, this work offers a methodological blueprint for bridging the gap between algorithmic theory and neural implementation.
Lallier, M.; Rius-Manau, C.; 23andMe Research Team, ; Carrion-Castillo, A.
Show abstract
Here, we test the hypothesis that early sustained exposure to complex bilingual environments can positively affect reading development by altering structural interhemispheric connectivity via the corpus callosum (CC). Interhemispheric connectivity has been shown to be inefficient in dyslexia, but also to support compensatory pathways when genetic risk for reading difficulties is present, by enabling the preserved right hemisphere to support a dysfunctional left hemisphere. Mediation models were conducted on children aged 9-10 years (with a 2-year follow-up assessment) from the Adolescent Brain Cognitive Development database (N>10,000). Polygenic scores (PGS) for dyslexia and cognitive performance and continuous bilingualism indices were used as predictors, with reading aloud as the outcome. Bilingualism showed a positive effect on reading partially mediated by the anterior CC, independently of overall brain size. In contrast, genetic predispositions to reading difficulties influenced reading primarily through overall brain size rather than CC connectivity specifically. These two pathways were independent, suggesting that bilingual experience and genetic risk operate through distinct neuroanatomical mechanisms. These findings suggest that recurrent early exposure to complex bilingual environments may shape the brains structural connectivity toward a more balanced and integrated bilateral frontal organisation. The results highlight potential brain compensatory pathways induced by environmental experiences that may support more efficient reading development and mitigate risks for developmental dyslexia.
Yao, J. K.; Mitchell, J.; Davison, A.; Yeatman, J. D.
Show abstract
Individual differences in cognitive abilities have been linked to variability in cortical folding, a stable neuroanatomical scaffold largely established in utero. In the domain of reading, recent findings in small groups of typical readers suggest that a sulcal interruption (superficial annectant gyrus, gyral gap) in the left posterior occipital temporal sulcus (lhpOTS) predicts better reading skills, posing the lhpOTS as a potential early biomarker of reading difficulties. However, it remains unknown whether this relationship found in typical readers generalizes to the dyslexic population and whether the lhpOTS can serve as a biomarker for dyslexia or predict response to targeted instruction.To fill these gaps, we examine the patterns of the lhpOTS in 209 children, including children with dyslexia, from four independently-collected samples. In typical readers, we find that the relationship between the lhpOTS and reading skills is robust, replicating across binary and continuous quantifications of the sulcal interruption. However, lhpOTS patterns neither distinguish dyslexic children from typical readers nor do they predict response to intervention. Instead, targeted reading intervention drives long-term gains in reading skills that are equivalent irrespective of VOTC anatomy. Together, these findings distinguish neuroanatomical correlates of skilled reading from determinants of reading impairment and learning capacity and emphasize the importance of the educational environment in supporting reading acquisition for children with dyslexia. SIGNIFICANCE STATEMENTEarly predictors of dyslexia are important for understanding the etiology of reading difficulties and informing early intervention. One candidate biomarker for dyslexia is the left posterior occipital temporal sulcus (lhpOTS), a neuroanatomical feature established before birth. In typical readers, the presence of an interruption in the lhpOTS has been linked to better reading skills. Here, we examine this neuroanatomical feature in 209 children with and without dyslexia. While the lhpOTS reliably relates to reading skill in typical readers, it neither differentiates dyslexic from typical readers nor predicts response to intensive reading intervention. These results show that brain anatomy reflects reading proficiency but does not constrain learning and highlights the power of targeted intervention to support reading development.
Coopmans, C. W.; de Hoop, H.; Tezcan, F.; Hagoort, P.; Martin, A. E.
Show abstract
Studies of perception have long shown that the brain adds information to its sensory analysis of the physical environment. A touchstone example for humans is language use: to comprehend a physical signal like speech, the brain must add linguistic knowledge, including syntax. Yet, syntactic rules and representations are atemporal (i.e., abstract and not bound by time), so they must be translated into time-varying signals for speech comprehension and production. Here, we test three different models of the temporal spell-out of syntactic structure against brain activity of people listening to Dutch stories: an integratory bottom-up parser, a predictive top-down parser, and a mildly predictive left-corner parser. These models build exactly the same structure but differ in when syntactic information is added by the brain - this difference is captured in the (temporal distribution of the) complexity metric incremental node count. Using temporal response function models with both acoustic and information-theoretic control predictors, node counts were regressed against source-reconstructed delta-band activity acquired with magnetoencephalography. Neural dynamics in left frontal and temporal regions most strongly reflect node counts derived by the top-down method, which postulates syntax early in time, suggesting that predictive structure building is an important component of Dutch sentence comprehension. The absence of strong effects of the left-corner model further suggests that its mildly predictive strategy does not represent Dutch language comprehension well, in contrast to what has been found for English. Understanding when the brain projects its knowledge of syntax onto speech, and whether this is done in language-specific ways, will inform and constrain the development of mechanistic models of syntactic-structure building in the brain.
de Heer Kloots, M.; Kazemian, A.; Turner, W.; Parvizi, J.; Gwilliams, L.
Show abstract
Context is critical for both human and artificial speech comprehension systems. While the role of preceding context in speech processing has been well documented, the neural mechanisms supporting the integration of subsequent input -- phonemes and words that occur in the future -- remain poorly understood. Here, we leverage advances in artificial speech systems to model the contribution of different sources of context on the neural encoding of speech in the human brain. For neural encoding, context-informed but not context-uninformed speech model embeddings explain unique variance in human neural activity beyond acoustics, including in early speech processing regions. In particular, model embeddings informed by past, future, and surrounding context explain activity in distinct intracranial electrodes. These electrodes are left-lateralised, and spatially intermixed in the temporal lobe. We find that beyond-word context is crucial for the representational quality of speech model embeddings, and in particular for the encoding of abstract linguistic information. Our finding that spatially neighboring yet distinct neural populations in the temporal lobe encode representations shaped by different contextual sources (past, future, and surrounding input) provides key insight into the neural circuitry that integrates multiple forms of contextual information. Furthermore, our results may inform the downstream use of self-supervised speech representations in language technology tasks, and in models of speech comprehension in the human brain.
Negi, A.; Oota, S. R.; Gupta, M.; Deniz, F.
Show abstract
Recent studies have demonstrated that fine-tuning language models with brain data can improve their semantic understanding, although these findings have so far been limited to English. Interestingly, similar to the shared multilingual embedding space of pretrained multilingual language models, human studies provide strong evidence for a shared semantic system in bilingual individuals. Here, we investigate whether fine-tuning language models with bilingual brain data changes model representations in a way that improves them across multiple languages. To test this, we fine-tune monolingual and multilingual language models using brain activity recorded while bilingual participants read stories in English and Chinese. We then evaluate how well these representations generalize to the bilingual participants first language, their second language, and several other languages that the participants are not fluent in. We assess the fine-tuned language models on brain encoding performance and downstream NLP tasks. Our results show that bilingual brain-informed fine-tuned language models outperform their vanilla (pretrained) counterparts in both brain encoding performance and most downstream NLP tasks across multiple languages. These findings suggest that brain-informed fine-tuning improves multilingual understanding in language models, offering a bridge between cognitive neuroscience and NLP research. We make our code publicly available. 2
Gillis, M.; Kries, J.; Wouters, J.; Gwilliams, L.; Vandermosten, M.
Show abstract
This study investigates the neural dynamics of phoneme processing in 7-year-old children with and without dyslexia (25;9 [male]), using EEG recordings collected during continuous speech listening. By applying temporal generalization to phonetic descriptor decoding, we can disentangle whether potential phoneme processing deficits are due to the maintenance of phonemes in verbal short-term memory and/or inferred differences in phonetic processing speed, both of which are thought to be impaired in dyslexia. We investigated whether phonetic processing depends on the phonemes position or its lexical competition. Our results reveal two key findings that may help explain the challenges faced by children with dyslexia. First, these children exhibit reduced decoding accuracy for word-onset phonemes, suggesting disruptions in either predictive, word-level anticipatory mechanisms or in the intrinsic rhythmic processing aligned with word boundaries. Second, they exhibit increased decoding accuracy for non-onset phonemes with low lexical competition approximately 400 ms after phoneme onset. This pattern suggests that children with dyslexia retain linguistically less relevant sounds longer in verbal short-term memory and process them more slowly compared to their typical reading peers. Together, these findings suggest that dyslexia is characterized by altered phonetic encoding strategies, specifically inefficient prioritization of relevant phonological information. This work provides new insight into the neural mechanisms underlying phonological deficits and contributes to a deeper understanding of the cognitive basis of dyslexia. Significance statementDyslexia is associated with difficulties in phonological processing. Investigating EEG during continuous speech listening, we show that children with dyslexia exhibit weaker encoding of word-onset phonemes and prolonged processing of less informative phonemes. These altered encoding strategies suggest inefficient prioritization of linguistic information, offering new insight into the neural basis of dyslexia.
Staples, R.; DeMarco, A. T.; Laks, A. B.; Turkeltaub, P. E.
Show abstract
Computational models are a linchpin in our understanding of the neurocognitive basis of reading. These models can simulate idealized profiles of alexia syndromes, but in reality, individuals with alexia present with a wide range of mixed deficits rather than idealized syndromes. To provide a complete cognitive theory of reading, computational models must be able to account for this individual variation. However, this has never been demonstrated. We test oral reading and non-reading phonological and semantic processing in 83 left-hemisphere stroke survivors. We show that individual alexia profiles can be simulated by applying graded phonology and semantic lesions to an artificial neural network model of reading, creating "matched models" that represent individual stroke survivors. The severity of damage to the semantic and phonological layers of the matched models was highly correlated with directly-measured semantic and phonological processing deficits. However, we also identify systematic ways in which the models fail to simulate the reading performance of their matched stroke survivors. Our results support theories of alexia that rely on process-based deficits, demonstrate the feasibility of large-scale individualized modelling of alexia, and suggest ways to further improve the correspondence of models and human reading behavior.
Regev, T. I.; Affourtit, J.; Chen, X.; Schipper, A. E.; Bergen, L.; Mahowald, K.; Fedorenko, E.
Show abstract
A network of left frontal and temporal brain regions supports high-level language processing-- including the processing of word meanings, as well as word-combinatorial processing--across presentation modalities. This core language network has been argued to store our knowledge of words and constructions as well as constraints on how those combine to form sentences. However, our linguistic knowledge additionally includes information about sounds (phonemes) and how they combine to form clusters, syllables, and words. Is this knowledge of phoneme combinatorics also represented in these language regions? Across five fMRI experiments, we investigated the sensitivity of high-level language processing brain regions to sub-lexical linguistic sound patterns by examining responses to diverse nonwords--sequences of sounds/letters that do not constitute real words (e.g., punes, silory, flope). We establish robust responses in the language network to visually (Experiment 1a, n=605) and auditorily (Experiments 1b, n=12, and 1c, n=13) presented nonwords relative to baseline. In Experiment 2 (n=16), we find stronger responses to nonwords that obey the phoneme-combinatorial constraints of English. Finally, in Experiment 3 (n=14) and a post-hoc analysis of Experiment 2, we provide suggestive evidence that the responses in Experiments 1 and 2 are not due to the activation of real words that share some phonology with the nonwords. The results suggest that knowledge of phoneme combinatorics and representations of sub-lexical linguistic sound patterns are stored within the same fronto-temporal network that stores higher-level linguistic knowledge and supports word and sentence comprehension.
Gao, C.; Ma, Z.; Chen, J.; Li, P.; Huang, S.; Li, J.
Show abstract
Transformer-based large language models (LLMs) have significantly advanced our understanding of meaning representation in the human brain. However, increasingly large LLMs have been questioned as valid cognitive models due to their extensive training data and their ability to access context thousands of words long. In this study, we investigated whether instruction tuning, another core technique in recent LLMs beyond mere scaling, can enhance models ability to capture linguistic information in the human brain. We evaluated the self-attention of base and fine-tuned LLMs of different sizes against human eye movement and functional magnetic resonance imaging (fMRI) activity patterns during naturalistic reading. We show that scaling has a greater impact than instruction tuning on model-brain alignment, reinforcing the scaling law in brain encoding performance. These finding have significant implications for understanding the cognitive plausibility of LLMs and their role in studying naturalistic language comprehension.
Li, J.; Luh, W.-M.; Pylkkanen, L.; Yang, Y.; Hale, J.
Show abstract
Human language processing involves not only combining word meanings in accordance with semantic and syntactic constraints, but also figuring out who and what is being referred to. Here we present a first study towards a mechanistic understanding of the neural basis for referential processing. Using both functional MRI and magnetoencephalography (MEG), we identified a consistent increase of activity in a network spanning the anterior and posterior left middle temporal gyrus and the angular gyrus for pronoun processing during naturalistic listening for both English and Chinese speakers. We then adopted a "reverse-engineering" approach to examine the cognitive processes underlying pronoun resolution. We evaluated the neural fit of three symbolic models that each formalizes a different strand of explanation for pronoun resolution in the cognitive and linguistic literature, as well as two deep neural network models with an LSTM or a Transformer architecture. Our results favor the memory-based symbolic model, suggesting a domain-general mechanism of pronoun resolution that resembles memory retrieval.
Balboni, I.; Kepinska, O.; Rampinini, A.; Berthele, R.; Golestani, N.
Show abstract
Understanding the cognitive architecture of the human language faculty requires exploring the boundaries of both predisposition and environmental experience. However, previous research on extraordinary multilingualism has often confounded language aptitude with multilingual experience, obscuring their distinct neural correlates. Here, we leveraged a linguistically diverse sample (N=121) and extensive behavioural testing to dissociate language aptitude from multilingual experience, modelling both dimensions continuously in whole-brain speech processing. Language aptitude and multilingual experience were weakly related, and their dissociation was also evident at the neural level. Higher language aptitude showed a neural signature of efficiency, characterised by lower activation in core perisylvian regions. In contrast, higher multilingualism was associated with greater engagement of regions implicated in narrative, multimodal, and memory processing, and with recruitment of traditional language hubs only during degraded speech processing, likely reflecting active attempts to decode unintelligible input. Finally, aptitude and experience interacted within sensorimotor regions. Continuous quantification of multilingual experience proved more sensitive than artificial grouping. By disentangling language aptitude from multilingual experience, this work provides a more precise account of the multilingual brain, and shows that its neurobiology can be better understood by modelling predisposition and experience as distinct but interacting dimensions.
Eden, G. F.; Coutinho, M. R.
Show abstract
Prior studies have reported inconsistent results for neuroanatomical differences between early bilinguals and monolinguals. These studies primarily measured gray matter volume (GMV), involved small samples, and prioritized adults. Few studies of early bilinguals have measured cortical thickness (CT), which offers more anatomical specificity. It remains unclear whether results derived from differing metrics and approaches (e.g., vertex-versus parcel-wise analyses) converge. Using data from the Adolescent Brain Cognitive DevelopmentSM (ABCD) Study, we compared neuroanatomy between large groups of early cultural Spanish-English bilingual and English monolingual children (N = 1,209) matched on age, pubertal status, sex, handedness, socioeconomic status (SES), and nonverbal reasoning. Whole-brain voxel-based morphometry revealed areas of greater and of lesser GMV in bilinguals than monolinguals across all lobes. Vertex-wise CT analyses similarly identified widespread differences, with bilinguals showing areas of both thicker and thinner cortex. We contextualized these findings with parcel-wise CT analyses (average CT values), utilizing two atlases of differing spatial granularity. Parcel-wise results showed good correspondence with vertex-wise findings when implementing the more fine-grained atlas (Destrieux), but use of the coarser atlas (Desikan-Killiany) provided results that led to different conclusions. Finally, we tested for interaction effects between bilingualism and SES on CT and found several regions where differences between bilinguals and monolinguals in CT were modulated by SES. Together, these findings indicate that early bilingualism is associated with extensive neuroanatomical differences relative to monolinguals during childhood, and that these results can vary as a function of neuroanatomical metric, analysis approach, atlas granularity, and SES. Research HighlightsEarly Spanish-English bilingual and monolingual children differ in gray matter volume and cortical thickness across multiple brain regions. Cortical thickness differences between bilinguals and monolinguals cannot be firmly attributed to adaptations associated with language or executive control. Socioeconomic status modulates cortical differences between early bilinguals and monolinguals, revealing unique thickness patterns for those with lower versus higher SES backgrounds. Parcel-wise between-group cortical thickness results are affected by atlas choice and can influence the interpretation of the findings.
Cara, C.; Zantonello, G.; Ghio, M.; Tettamanti, M.
Show abstract
Dyslexia is a neurobiological disorder characterised by reading difficulties, yet its underlying causes remain unclear. Neuroimaging and behavioural studies found anomalous responses in tasks requiring phonological processing, motion perception, and implicit learning, and showed gray and white matter abnormalities in several brain regions of dyslexics compared to controls, indicating that dyslexia is a heterogeneous condition and promoting a multifactorial approach. In order to evaluate whether the combination of behavioural and multimodal MRI can have greater sensitivity in identifying neurocognitive traits of dyslexia compared to monocomponential approaches, in 19 dyslexic and 19 control subjects we acquired behavioural cognitive assessments, multiple (phonological, visual motion, rhythmic) mismatch-response functional MRI tasks, structural diffusion-weighted and T1-weighted images. To examine between-group differences in the multimodal neurocognitive measures, we applied univariate and multivariate approaches. Results showed that dyslexics performed worse than controls in behavioural phonological tasks. Neuroimaging analyses revealed that individuals with dyslexia present reduced cerebellar responses to mismatching rhythmic stimuli, as well as structural disorganization in several white matter tracts and cortical regions previously implicated in dyslexia. Most importantly, in line with the view of dyslexia as a multifactorial phenomenon, a machine learning model trained with features from all three MRI modalities (functional, diffusion, and T1-weighted) discriminated between dyslexics and controls with greater accuracy than models including just one modality. The individual classification scores in the multimodal machine learning model correlated with behavioural reading accuracy. These results confirm that dyslexia should be approached as a composite condition characterised by multiple distinctive cognitive and brain features.
Chen, X. J.; Blanco-Elorrieta, E.
Show abstract
Bilingualism research has long been challenged by a lack of a unified approach to quantifying language dominance and degree of multilingualism. While numerous questionnaires (e.g., LHQ, BLP, LEAP-Q, and LUQ) provide valuable data on language background variables, they lack a standardized formula to compute key measures from it. We introduce two formulas that synthesize critical linguistic variables to efficiently calculate language dominance and a multilingualism score that ranges from perfect monolingualism to native-like proficiency in multiple languages. Validation across two large datasets shows our dominance measure closely aligns with more complex PCA methods while being simpler and more efficient.
Puffay, C.; Vanthornhout, J.; Gillis, M.; De Clercq, P.; Accou, B.; Van hamme, H.; Francart, T.
Show abstract
When a person listens to natural speech, the relation between features of the speech signal and the corresponding evoked electroencephalogram (EEG) is indicative of neural processing of the speech signal. Using linguistic representations of speech, we investigate the differences in neural processing between speech in a native and foreign language that is not understood. We conducted experiments using three stimuli: a comprehensible language, an incomprehensible language, and randomly shuffled words from a comprehensible language, while recording the EEG signal of native Dutch-speaking participants. We modeled the neural tracking of linguistic features of the speech signals using a deep-learning model in a match-mismatch task that relates EEG signals to speech, while accounting for lexical segmentation features reflecting acoustic processing. The deep learning model effectively classifies languages. We also observed significant differences in tracking patterns between comprehensible and incomprehensible speech stimuli within the same language. It demonstrates the potential of deep learning frameworks in measuring speech understanding objectively.
Chacon, D. A.; Shrestha, S.; Dillon, B. W.; Bhatt, R.; Almeida, D.; Marantz, A.
Show abstract
At first glance, the brains language network appears to be universal, but languages clearly differ. How does the brain adapt to the specific details of individual grammatical systems? Here, we present an MEG study on case and agreement in Hindi and Nepali. Both languages use split-ergative case systems. However, these systems interact with verb agreement differently - in Hindi, case features conspire to determine which noun phrase (NP) the verb agrees with (subject, object, or neither), but in Nepali the verb always agrees with the subject NP. We found that left inferior frontal and left anterior temporal regions are sensitive to case features in both languages. Across case configurations, these same brain areas in Hindi participants show different patterns of activity for sentences that require masculine vs. feminine marking on the verb, before the comprehenders encounter it. Additionally, the left temporoparietal junction in Hindi shows different activity for subject and object agreement configurations. Both findings are not observed in Nepali participants. We suggest that this brain response demonstrates a unique-to-Hindi selection of an agreement controller and pre-encoding of the verbs morphological features. This shows that brain activity reflects psycholinguistic processes that are intimately tied to grammatical features. HighlightsO_LIThe left inferior frontal lobe and the left anterior temporal lobe distinguish accusative objects versus bare object NPs in Hindi and Nepali, and pre-emptively encode gender agreement features in Hindi. C_LIO_LIThe left inferior parietal lobe shows a differential sensitivity to object-agreement and subject-agreement constructions in Hindi that is absent in Nepali C_LIO_LIMEG can reveal differences in neural activity that reflect specific requirements of different grammatical systems C_LI
Ahamdi, S. S.; Fridriksson, J.; Den Ouden, D.
Show abstract
Language impairments in aphasia are characterized by various representational disruptions that may be reflected in discourse production. This research examines the capacity of transformer-based language models, particularly GPT-2, to serve as a computational framework for analyzing variations in aphasic narrative speech. A longitudinal dataset of narrative speech samples collected at six time points from individuals with aphasia (N = 47) was utilized as part of an intervention study. All transcripts were processed via the GPT-2 language model to obtain activation values from each of the 12 transformer layers. Statistically significant differences in activation magnitude across aphasia subtypes were found at every layer (all p < .001), with the most pronounced effects in the deeper layers. Pairwise Tukey HSD tests revealed consistent distinctions between Brocas aphasia and both Anomic and Wernickes aphasia, suggesting a shared activation profile between the latter two. Longitudinal tests revealed significant changes over time, especially in the final three layers (10-12). These findings suggest that transformer-based activation patterns reflect meaningful variation in aphasic discourse and could complement current diagnostic tools. Overall, GPT-2 provides a scalable tool to model representational dynamics in aphasia and enhance the clinical interpretability of deep language models.