Back

A syntax network morphospace reveals sudden transitions between discrete developmental stages during language acquisition

Barcelo-Coblijn, L.; Eguiluz, V. M.; Seoane, L. F.

2025-02-19 systems biology
10.1101/2025.02.19.639060 bioRxiv
Show abstract

Syntax is an aspect of human language responsible for the hierarchical ordering of linguistic structures. Syntax can be summarized by dependency trees with words as nodes and edges reflecting syntactic subordination. By merging trees from several sentences, we obtain syntax graphs or networks, which display distinct shapes depending on whether the language capability is well-formed, still developing, or pathological. Such graphs make syntactic capacity quantifiable at a systemic level, revealing emerging patterns and universalities. What is the structure of syntax networks during ontogeny in typically developing (TD) children? Do cognitively challenged children develop language through alternative routes? Here we quantify and portray the typical development of syntax networks in Dutch, and find that children affected by Down syndrome, hearing impairment, and specific language impairment initially seem to follow the typical developmental path but eventually halt, culminating in a different linguistic phenotype. Our expanded data set (with almost 50 times more data than earlier studies) and increased mathematical dimensions to quantify network shape enable us to: (i) confirm and refine a proposed sharp transition in language development, (ii) correlate specific network traits with syntax maturation, and (iii) quantify the aspects that fall short in atypical development--suggesting potential diagnostic tools. We also find grounds to hypothesize a gap in syntax maturation, separating challenged children who nevertheless reach the latest stage from others systematically stuck, regardless of their condition. Our quantitative analysis enables a rigorous visualization of linguistic development trajectories, an old (yet mostly qualitative) theme in linguistics. Similar works should allow to test and propose specific hypotheses on solid grounds, as we do here. Future efforts should generalize to other languages and/or clinical conditions, seeking patterns that might point at universalities in language development.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.