Back

Journal of Speech, Language, and Hearing Research

American Speech Language Hearing Association

All preprints, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
A Meta-Analytical Review of Executive Function Skills in Adults who Stutter

Ofoe, L. C.; Ntourou, K.; Clifton, S.; Coalson, G. A.

2025-09-04 pathology 10.1101/2025.09.02.25334917 medRxiv
Top 0.1%
63.2%
Show abstract

PurposeExecutive function has been identified as a potential area of vulnerability in individuals who stutter. The present study identified and analyzed the data across empirical studies of the executive function skills of adults who do (AWS) and do not stutter (AWNS). MethodElectronic databases, literature reviews, and reference sections of articles and dissertations were searched to identify candidate studies that examined behavioral measures of working memory, inhibition, and/or cognitive flexibility. A total of 39 studies met the eligibility criteria for this meta-analysis. A random-effects model was applied to estimate the pooled effect sizes (Hedges g) and 95% confidence intervals. ResultsAWS were significantly less accurate than AWNS on measures of working memory (Hedges g = - 0.41, p < .001), including nonword repetition (Hedges g = -.57, p < .001), forward digit span (Hedges g = -0.23, p = .02), backward digit span (Hedges g = -.38, p = .004) and operation span tasks (Hedges g = - .37, p = .018). AWS performed comparably to AWNS on inhibition measures (Hedges g = -0.10, p = .37). An insufficient number of published studies were available to conduct a meaningful analysis of cognitive flexibility. ConclusionsPresent findings suggest that AWS, as a group, exhibit weaknesses in one component of executive function - working memory - compared to AWNS. Additional research is necessary to determine potential differences in inhibition and cognitive flexibility in AWS.

2
Late-Talking Children Talk More? A Machine Learning Approach to Speech Act Analysis in Early Childhood

Dhakal, G.; He, H.; Newman, S. D.; Xiong, Y.

2025-10-02 pathology 10.1101/2025.09.30.679667 medRxiv
Top 0.1%
47.4%
Show abstract

Speech acts shape early language development and social cognition, yet little is known about how late-talking (LT) children use them to achieve communicative goals. We compared LT and typically developing (TD) preschoolers (1;09-6;00) across nine dyadic English corpora, using a Conditional Random Field model to annotate speech acts. We analyzed speech act distributions, hierarchical relations, and contingent responses to assess production and comprehension. TD children produced more declarative statements and wh-questions, whereas LT children produced more unclear word-like utterances and showed reduced comprehension ability (1;09-2;07). Speech acts classified LT and TD groups with 72.3% accuracy, improving to 76.6% with linguistic and demographic features. Classification was driven by co-occurring patterns of speech act frequencies. LT children showed delayed onset of speech acts but employed more speech acts after 3;09, focusing on speaker-centered goals, whereas TD children favored collaborative use of speech acts, revealing complex dynamics in the development of communicative skills.

3
Self-reported effects of classic psychedelics on stuttering

Gold, N.; Goldway, N.; Gerlach-Houck, H.; Jackson, E. S.

2023-04-20 pathology 10.1101/2023.04.18.537312 medRxiv
Top 0.1%
25.6%
Show abstract

Stuttering is a neurodevelopmental communication disorder that can lead to significant social, occupational, and educational challenges. Traditional behavioral interventions for stuttering can be helpful, but effects are often limited. Classic psychedelics hold promise as a complement to traditional interventions, but their impact on stuttering is unknown. We conducted a qualitative content analysis to explore potential benefits and negative effects of psychedelics on stuttering using publicly available Reddit posts. A combined inductive-deductive approach was used whereby meaningful units were extracted and codes were initially assigned inductively. We then deductively applied an established framework to organize the effects (i.e., codes) into five subthemes (Behavioral, Emotional, Cognitive, Belief, and Social Connection), each of which was grouped under an organizing theme (positive, negative, neutral). Results indicated that the effects of psychedelics spanned all subthemes. Nearly 75% of participants reported overall positive effects. Nearly 60% of participants indicated positive behavioral change (e.g., reduced stuttering, increased speech control), 40% reported positive emotional benefit, 15% reported positive cognitive changes, 12% reported positive effects on beliefs, and 7% indicated positive social effects. Approximately 10% of participants reported negative behavioral effects (e.g., increased stuttering, reduced speech control). Psychedelics may help many stutterers improve communication, cultivate a healthier outlook, and promote psychological well-being. These preliminary results indicate that future clinical trials investigating psychedelic-assisted speech therapy for stuttering are warranted.

4
EEG responses to auditory cues predict fluency variability and stuttering intervention outcome

Rocha, M. F.; Carmona, J.; Correia, J. M.

2025-02-24 neuroscience 10.1101/2025.02.21.635719 medRxiv
Top 0.1%
18.7%
Show abstract

Stuttering is a variable speech disorder whose brain mechanisms remain unknown. Sensorimotor brain circuits, critical for motor-speech control, including auditory processing necessary for speech prediction and monitoring, have been linked to the disorder. Despite considerable advances, it remains unclear whether auditory circuits relate to stuttering variability, and whether the panoply of interventions for persons who stutter can lead to brain changes within these circuits. We employed electroencephalography (EEG), in a group of persons who stutter, in combination with auditory probes to tap onto the importance of auditory cortical regions in stuttering variability. Participantsproduced flexible speech (i.e., describing visual scenes) and non-flexible speech (i.e., reading syllables), following an auditory cue. More pronounced P200 auditory evoked potentials were observed in participants with higher dysfluency rates, mainly in the spontaneous speech task. Interestingly, speech therapy intervention led to a reduction of the P200 potential, which was in turn significantly related to fluency improvements. Furthermore, EEG response patterns discriminative of cue frequency (400 or 800 Hz tones) were also predictive of dysfluency scores. Our study highlights the involvement of auditory cortical processing and that of auditory attention in stuttering variability. We support that a higher state of auditory alertness may be implicated in the sensorimotor mechanisms of stuttering, and that speech therapy interventions promoting more self-confident communication can restraint auditory alertness, and potentially reduce speech dysfluencies. HighlightsO_LIAuditory probes can assess the auditory cortex in speech production and stuttering. C_LIO_LIStuttering severity correlates to EEG auditory responses during speech preparation. C_LIO_LIHigher states of auditory alertness in stuttering may be reduced by speech therapy. C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=123 SRC="FIGDIR/small/635719v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@18a639dorg.highwire.dtl.DTLVardef@91efc6org.highwire.dtl.DTLVardef@114dd97org.highwire.dtl.DTLVardef@e01024_HPS_FORMAT_FIGEXP M_FIG C_FIG Speaking requires orchestrating several brain processes at a time. The auditory system assumes a central role, not only in waiting for the right moment to initiate speech, listening to self-produced speech, predicting the consequence of future speech, but also adjusting these processes to the intermittent nature of stuttering.

5
Psychiatric Voice Biomarkers: Methodological flaws in pediatric populations

Hamoudi, H. J. A. S.; Wu, M.-J.; Sanches, M.; Soutullo, C. A.; Olmos, C.; Taylor, L. K.; Zunta-Soares, G.; Soares, J. C.; Mwangi, B.

2025-10-15 psychiatry and clinical psychology 10.1101/2025.10.13.25337901 medRxiv
Top 0.1%
15.7%
Show abstract

IntroductionPsychiatric assessments rely on patient self-reports, clinician observations, and standardized scales, while objective technological tools are currently not reliable enough to be utilized in a clinical setting. Voice may be utilized as a biomarker in different scenarios, including differential diagnosis, assessing symptom severity and predicting suicidality. However, its use depends on accurate automatic speech recognition (ASR). Current gold standard open source ASR systems are trained mainly on adult speech and perform poorly in children, limiting application in pediatric psychiatry. MethodsWe benchmarked two open-source ASR models--NVIDIA Parakeet and Whisper-small--on the Ohio Child Speech Corpus (303 children, ages 4-9), using the reference human transcripts provided with the dataset. Audio was standardized to each models expected sampling rate. No model fine-tuning or adaptation was performed. For each utterance, we computed word error rate (WER) and character error rate (CER), and assessed semantic fidelity using Sentence Movers Distance (SMD) and BERTScore F1. Metrics were summarized overall, stratified by single-year age bins (4, 5, 6, 7, 8, 9), and also grouped into two broader categories: younger children (ages 4-6) and older children (ages 7-9). We compared WER, CER, SMD, and BERTScore F1 across both age groups and evaluated age effects as trends using nonparametric statistical tests. ResultsBoth models showed significant age effects where younger children had markedly higher word error rates (WER >40%) and character error rates (CER >30%) compared to older children (WER [~]30%, CER [~]20%). Sentence mover distance improved with age, while BERTScore F1 remained stable. Despite age-related improvements, overall transcription accuracy was low. DiscussionCurrent commonly used open-source ASR systems are inadequate for pediatric audio transcription, specifically in younger children. In order to build clinically translatable tools, collecting child-specific data and model fine-tuning through structured speech paradigms is essential.

6
Does over-reliance on auditory feedback cause disfluency? An fMRI study of induced fluency in people who stutter.

Meekings, S.; Jasmin, K.; Lima, C. F.; Scott, S. K.

2020-11-20 neuroscience 10.1101/2020.11.18.378265 medRxiv
Top 0.1%
15.1%
Show abstract

This study tested the idea that stuttering is caused by over-reliance on auditory feedback. The theory is motivated by the observation that many fluency-inducing situations, such as synchronised speech and masked speech, alter or obscure the talkers feedback. Typical speakers show speaking-induced suppression of neural activation in superior temporal gyrus (STG) during self-produced vocalisation, compared to listening to recorded speech. If people who stutter over-attend to auditory feedback, they may lack this suppression response. In a 1.5T fMRI scanner, people who stutter spoke in synchrony with an experimenter, in synchrony with a recording, on their own, in noise, listened to the experimenter speaking and read silently. Behavioural testing outside the scanner demonstrated that synchronising with another talker resulted in a marked increase in fluency regardless of baseline stuttering severity. In the scanner, participants stuttered most when they spoke alone, and least when they synchronised with a live talker. There was no reduction in STG activity in the Speak Alone condition, when participants stuttered most. There was also strong activity in STG in response to the two synchronised speech conditions, when participants stuttered least, suggesting that either stuttering does not result from over-reliance on feedback, or that the STG activation seen here does not reflect speech feedback monitoring. We discuss this result with reference to neural responses seen in the typical population.

7
Automated Macrolinguistic Discourse Analysis for Transdiagnostic Detection of Language Impairments

Lee, S. H.; Wang, S.; Varkanitsa, M.; Kiran, S.

2026-05-21 neurology 10.64898/2026.05.19.26353614 medRxiv
Top 0.1%
13.0%
Show abstract

Macrolinguistic discourse analysis offers valuable insight into how patients with neurogenic communication disorders organize and produce informative speech, yet it remains a largely manual and labor-intensive process. We report an automated pipeline for macrolinguistic discourse analysis for individuals with aphasia and dementia that integrates automatic speech recognition (ASR), utterance segmentation, sentence-level embeddings, centroid-based main-concept matching, and rule-based coherence error classification. These algorithms were applied to Cinderella story retellings from 309 participants (113 controls, 102 post-stroke aphasia (PWA), and 94 dementia). The algorithm reliably identified main concepts (83% accuracy against human labels) and derived interpretable features such as semantic distance to a main concept centroid, main concept coverage, and coherence error rates. Crucially, diagnostic classification results showed that logistic-regression classifiers trained on 10 macrolinguistic features distinguished aphasia from controls with high accuracy (AUC {approx} 0.94) but showed weaker separation for dementia (controls vs dementia AUC {approx} 0.66; aphasia vs dementia AUC {approx} 0.58). Semantic distance to the centroid emerged as a robust, informative predictor for diagnostic classification, demonstrating that the ability to produce narrative-aligned speech is clinically important. The automated pipeline enables scalable macrolinguistic discourse analysis that could support screening and longitudinal monitoring of discourse impairments across neurogenic populations.

8
Functional Roles of Sensorimotor Alpha and Beta Oscillations in Overt Speech Production

Huang, L. Z.; Cao, Y.; Janse, E.; Piai, V.

2024-10-08 neuroscience 10.1101/2024.09.04.611312 medRxiv
Top 0.1%
12.4%
Show abstract

Power decreases, or desynchronization, of sensorimotor alpha and beta oscillations (i.e., alpha and beta ERD) have long been considered as indices of sensorimotor control in overt speech production. However, their specific functional roles are not well understood. Hence, we first conducted a systematic review to investigate how these two oscillations are modulated by speech motor tasks in typically fluent speakers (TFS) and in persons who stutter (PWS). Eleven EEG/MEG papers with source localization were included in our systematic review. The results revealed consistent alpha and beta ERD in the sensorimotor cortex of TFS and PWS. Furthermore, the results suggested that sensorimotor alpha and beta ERD may be functionally dissociable, with alpha related to (somato-)sensory feedback processing during articulation and beta related to motor processes throughout planning and articulation. To (partly) test this hypothesis of a potential functional dissociation between alpha and beta ERD, we then analyzed existing intracranial electroencephalography (iEEG) data from the primary somatosensory cortex (S1) of picture naming. We found moderate evidence for alpha, but not beta, ERDs sensitivity to speech movements in S1, lending supporting evidence for the functional dissociation hypothesis identified by the systematic review.

9
Phonemic awareness deficits in an alphasyllabary language: Effects of task type and linguistic complexity in children with Specific Learning Disorder-Reading

Soman, A.; Dev, S. S.; Ravindren, R.

2026-04-07 psychiatry and clinical psychology 10.64898/2026.04.02.26349894 medRxiv
Top 0.1%
12.3%
Show abstract

Background Phonemic awareness deficits are a core feature of Specific Learning Disorder-Reading (SLD-R). How task- and language-specific factors influence these deficits in alphasyllabary languages may help clarify the cognitive mechanisms underlying reading impairment in SLD-R. Methods Thirty children with a DSM-5 diagnosis of SLD-R (mean age 11.4 years) and 29 age-matched typically developing children were given phoneme blending (words and pseudowords) and segmentation tasks in Malayalam. The effects of age and consonant clusters on task performance were evaluated. Results Children with SLD-R performed significantly worse than controls across most phonemic awareness tasks, with the largest deficits observed in pseudoword blending and word blending, and smaller deficits in segmentation. No significant difference was observed for initial phoneme deletion. In typically developing children, age showed strong positive correlations with phonemic performance across most tasks, whereas the SLD-R group showed weak or absent correlations, except in word blending and initial phoneme deletion. Consonant clusters significantly affected performance in both groups, with SLD-R showing more severe deficits. Conclusions Phonemic awareness deficits observed in SLD-R in alphasyllabary languages like Malayalam are more prominent in tasks where lexical support is absent, like pseudoword blending. These deficits vary across task types and linguistic complexity. Phonemic awareness improves with age in typically developing children, while improvement is uneven in children with SLD-R. The findings suggest that phonemic awareness deficits are a core feature of SLD-R across languages, but their manifestation is shaped by orthographic and linguistic characteristics of the writing system.

10
British Version of the Iowa Test of Consonant Perception

Guo, X.; Benzaquen, E.; Holmes, E.; Choi, I.; McMurray, B.; Bamiou, D.-E.; Berger, J. I.; Griffiths, T. D.

2024-09-07 neuroscience 10.1101/2024.09.04.611204 medRxiv
Top 0.1%
12.3%
Show abstract

The Iowa Test of Consonant Perception (ITCP) is a single-word closed-set speech- in-noise test with well-balanced phonetic features that provides a reliable testing option for real-world listening. Objectives. The current study aimed to establish a UK version of the test (B-ITCP) based on the British received pronunciation. Design. We conducted a validity test with 46 participants using the B-ITCP test, a sentence-in- noise test, and audiogram. Results. The B-ITCP demonstrated excellent test-retest reliability, cross-talker validity, and good convergent validity, consistent with the US results. Conclusions. These findings suggest that B-ITCP is a reliable measure of speech-in-noise perception, to facilitate comparative or combined studies in USA and UK. All materials (application and scripts) to run or construct the B-ITCP and ITCP are freely available online.

11
Altered Speech Processing in Childhood Listening Difficulties as Revealed by Chirped Speech Event-Related Potentials

Petley, L.; Wicks, T.; Miller, L. M.; Blankenship, C.; Chatwin, J.; Bormann, B. M.; Whittle, R. S.; Moore, D. R.

2026-08-17 otolaryngology 10.64898/2026.08.13.26360392 medRxiv
Top 0.1%
12.1%
Show abstract

Objective: Impaired understanding of noisy or degraded speech is a central feature of listening difficulties (LiD), but the possible causes of these symptoms are wide-ranging. Accordingly, recent research underscores the need to study these deficits using a test battery approach. Event-related potentials are useful objective metrics for studying LiD, but probing function across the speech processing hierarchy using traditional protocols is sequential and unrealistic in clinical settings. The novel chirped speech (Cheech) method combines natural speech with acoustic chirps to overcome these limitations. This study examines its utility for profiling childhood LiD. Methods: Twenty-eight children (15 typically developing, 13 with LiD), aged 8-17 years old, listened to a 17-minute Cheech story and detected a target word within the story via button press while EEG data were collected from 53 scalp sites. Results: Cheech successfully evoked responses from the auditory brainstem response through to the brain's language centers, as reflected by the N400 effect. Unlike TD children, those with LiD demonstrated N400 effects with atypical distributions that favored frontal rather than the typical parietal sites. A trend towards a delayed and reduced amplitude Wave V was also observed. Conclusions: Hierarchical examination of speech processing using Cheech primarily implicates altered language processing as a contributing factor to LiD, with the frontal topography of the N400 effect for those with LiD potentially suggesting a greater reliance on deliberate memory retrieval during the speech perception task. Significance: LiD could arise due to auditory and/or cognitive factors. The present results demonstrate the feasibility of objective, parallel measurement across this hierarchy and point to impaired language processing as a possible mechanism.

12
The impact of cognitive ability on multitalker speech perception in neurodivergent individuals

Lau, B. K.; Emmons, K.; Maddox, R. K.; Estes, A.; Dager, S.; (Astley) Hemingway, S.; Lee, A. K.

2022-09-20 psychiatry and clinical psychology 10.1101/2022.09.19.22280007 medRxiv
Top 0.1%
12.1%
Show abstract

The ability to selectively attend to one talker in the presence of competing talkers is crucial to communication. Here we investigate whether cognitive deficits in the absences of hearing loss can impair speech perception. We tested typical hearing, neurodivergent adolescents/adults with autism spectrum disorder, fetal alcohol spectrum disorder, and an age- and sex-matched neurotypical group. We found a strong correlation between IQ and speech perception, with individuals with lower IQ scores having worse speech thresholds. These results demonstrate that deficits in cognitive ability, despite intact peripheral encoding, can impair listening under complex conditions. These findings have important implications for conceptual models of speech perception and for audiological services to improve communication in real-world environments for neurodivergent individuals.

13
Cross-Linguistic Analysis of Speech Markers: Insights from English, Chinese, and Italian Speakers

Santi, G. C.; Catricala, E.; Kwan, S.; Wong, A.; Ezzes, Z.; Wauters, L.; Esposito, V.; Conca, F.; Gibbons, D.; Fernandez, E.; Santos-Santos, M. A.; Chen, T.-F.; Kwan-Chen, L. L.-Y.; Lo, R. R.; Tsoh, J.; Lung-Tat Chen, A.; Garcia, A. M.; de Leon, J.; Miller, Z.; Vonk, J. M. J.; Bruffaerts, R.; Grasso, S. M.; Allen, I. E.; Cappa, S. F.; Gorno-Tempini, M.-L.; Tee, B. L.

2024-10-16 neurology 10.1101/2024.10.15.24314191 medRxiv
Top 0.1%
12.0%
Show abstract

Cross-linguistic studies with healthy individuals are vital, as they can reveal typologically common and different patterns while providing tailored benchmarks for patient studies. Nevertheless, cross-linguistic differences in narrative speech production, particularly among speakers of languages belonging to distinct language families, have been inadequately investigated. Using a picture description task, we analyze cross-linguistic variations in connected speech production across three linguistically diverse groups of cognitively normal participants--English, Chinese (Mandarin and Cantonese), and Italian speakers. We extracted 28 linguistic features, encompassing phonological, lexico-semantic, morpho-syntactic, and discourse/pragmatic domains. We utilized a semi-automated approach with Computerized Language ANalysis (CLAN) to compare the frequency of production of various linguistic features across the three language groups. Our findings revealed distinct proportional differences in linguistic feature usage among English, Chinese, and Italian speakers. Specifically, we found a reduced production of prepositions, conjunctions, and pronouns, and increased adverb use in the Chinese-speakers compared to the other two languages. Furthermore, English participants produced a higher proportion of prepositions, while Italian speakers produced significantly more conjunctions and empty pauses than the other groups. These findings demonstrate that the frequency of specific linguistic phenomena varies across languages, even when using the same harmonized task. This underscores the critical need to develop linguistically tailored language assessment tools and to identify speech markers that are appropriate for aphasia patients across different languages.

14
Semantic and phonetic markers in schizophrenia-spectrum disorders; a combinatory machine learning approach

Voppel, A.; de Boer, J.; Brederoo, S.; Schnack, H.; Sommer, I. e. c.

2022-07-15 psychiatry and clinical psychology 10.1101/2022.07.13.22277577 medRxiv
Top 0.1%
11.9%
Show abstract

IntroductionSpeech is a promising marker for schizophrenia-spectrum disorder diagnosis, as it closely reflects symptoms. Previous approaches have made use of different feature domains of speech in classification, including semantic and phonetic features. However, an examination of the relative contribution and accuracy per domain remains an area of active investigation. Here, we examine these domains (i.e. phonetic and semantic) separately and in combination. MethodsUsing a semi-structured interview with neutral topics, speech of 94 schizophrenia-spectrum subjects (SSD) and 73 healthy controls (HC) was recorded. Phonetic features were extracted using a standardized feature set, and transcribed interviews were used to assess word connectedness using a word2vec model. Separate cross-validated random forest classifiers were trained on each feature domain. A third, combinatory classifier was used to combine features from both domains. ResultsThe phonetic domain random forest achieved 81% accuracy in classifying SSD from HC. For the semantic domain, the classifier reached an accuracy of 80% with a sparse set of features with 10-fold cross-validation. Joining features from the domains, the combined classifier reached 85% accuracy, significantly improving on models trained on separate domains. Top features were fragmented speech for phonetic and variance of connectedness for semantic, with both being the top features for the combined classifier. DiscussionBoth semantic and phonetic domains achieved similar results compared with previous research. Combining these features shows the relative value of each domain, as well as the increased classification performance from implementing features from multiple domains. Explainability of models and their feature importance is a requirement for future clinical applications.

15
Automated Phonological Error Scoring for Children with Language and Hearing Impairment

Sundstrom, S.; Themistocleous, C.

2024-09-04 neurology 10.1101/2024.09.04.24313011 medRxiv
Top 0.1%
11.9%
Show abstract

PurposePhonological production impairments are prevalent in children with developmental language disorder (DLD) and hearing impairment (HI). This study aims to quantify and compare phonological errors in Swedish-speaking children using a novel automated assessment tool and provide an automatic machine learning classification algorithm of children with DLD and HI to age-matched controls based on phonological errors. Methods72 Swedish-speaking children (29 with DLD, 14 with HI, and 29 typically developing) participated. Phonological production was elicited using a 72-item confrontation naming task. A novel tool was developed to calculate a composite phonological error score and specific scores for different phonological errors (deletions, insertions, substitutions, and transpositions) from written speech productions. This tool leverages the International Phonetic Alphabet (IPA) and a form of the Normalized Damerau-Levenshtein Distance for accurate error analysis. ResultsThe composite score successfully differentiated between children with DLD and typically developing children, highlighting its sensitivity in detecting phonological impairment. Machine learning models can accurately differentiate between children with and without language disorders. However, children with DLD and HI differed in the phonemic deletion errors, which suggests that their phonemic production is relatively similar. ConclusionsChildren with DLD and HI exhibit significantly higher phonological error rates compared to typically developing peers. Children with HI and DLC are comparably impaired in phonology (as manifested by the compositive phonological score). These findings highlight the potential of machine learning for early identification and targeted intervention in language disorders, improving outcomes for affected children and demonstrated the potential of a multilingual tool for scoring phonological errors.

16
Effect of Spatial Release from Masking on Listening Effort in Different Semantic Contexts

Dantanarayana, N. D.; Li, Y.; Litovsky, R. Y.; Borjigin, A.

2026-07-20 neuroscience 10.64898/2026.07.13.738303 medRxiv
Top 0.1%
11.6%
Show abstract

Humans often communicate and learn in noisy, complex listening environments. Here, we investigated the effects of spatial hearing and semantic context cues on speech intelligibility and listening effort in young adults with typical hearing. The listening task included conditions in which target speech and speech maskers were either spatially co-located or separated. Target sentences were either semantically coherent or anomalous, while the masker comprised a mixture of two coherent sentences. Results showed higher speech intelligibility in spatially separated than co-located conditions, demonstrating a robust spatial release from masking (SRM), which is consistent with prior findings. SRM did not differ between semantically coherent and anomalous sentences, indicating comparable benefits of spatial cues across semantic contexts. However, within each spatial configuration, intelligibility was higher for coherent than anomalous sentences. Listening effort, indexed by peak pupil dilation in pupillometry measurement, was reduced in spatially separated conditions, suggesting a trend toward a release from listening effort. Analysis of the timing of peak pupil dilation revealed a significantly delayed peak dilation for anomalous sentences in the co-located condition compared with coherent sentences in the separated condition, indicating increased processing demands in the absence of spatial and semantic cues. Finally, SRM was correlated with the magnitude of release from listening effort for coherent sentences, but not for anomalous sentences, suggesting that intelligibility and listening effort benefits might co-occur when contextual cues are available.

17
Automated detection of adult autism from vowel acoustics using machine learning

Georgiou, G. P.; Paphiti, M.

2026-04-04 health informatics 10.64898/2026.04.03.26350102 medRxiv
Top 0.1%
11.3%
Show abstract

Autism spectrum disorder (ASD) is a neurodevelopmental condition for which timely and accurate detection remains a major clinical priority. Early and reliable identification is important because it can facilitate access to assessment, diagnosis, and appropriate support; however, current diagnostic pathways still rely largely on behavioural evaluation and clinical judgement. In this context, machine-learning (ML) approaches have attracted growing interest because they can identify subtle and complex patterns in speech data that may not be easily captured through conventional methods. The current study capitalizes on this potential by developing and evaluating ML models for distinguishing autistic individuals from neurotypical individuals based on speech features. More specifically, acoustic features of vowels, including fundamental frequency (F0), first three formants (F1, F2, F3), duration, jitter, shimmer, harmonics-to-noise ratio (HNR), and intensity, were elicited from 18 autistic adults and 18 neurotypical adults through a controlled production task. Then, four supervised ML models were trained and evaluated on these features: LightGBM, Random Forest, Support Vector Machine, and XGBoost. All models demonstrated good classification performance, with the best-performing model achieving a strong discriminability of 89%. The explainability analysis identified F0 as the most influential predictor by a substantial margin, followed by intensity, F3, and F1, while duration, shimmer, HNR, jitter, and F2 contributed more modestly. These findings demonstrate that vowel acoustics contain clinically relevant information for distinguishing autistic from neurotypical adult speech and highlight the potential of interpretable, speech-based ML as a transparent and scalable aid for ASD screening and assessment.

18
Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations

Muller, B.; Ortiz Barranon, A. A.; Roberts, L.

2026-04-17 neurology 10.64898/2026.04.12.26350731 medRxiv
Top 0.1%
11.2%
Show abstract

Dysarthric speech severity assessment typically requires either trained clinicians or supervised machine learning models built from labelled pathological speech data, limiting scalability across languages and clinical settings. We present a training-free method (no supervised severity model is trained; feature directions are estimated from healthy control speech using a pretrained forced aligner) that quantifies dysarthria severity by measuring the degradation of phonological feature subspaces within frozen HuBERT representations. For each speaker, we extract phone-level embeddings via Montreal Forced Aligner, compute d scores along phonological contrast directions (nasality, voicing, stridency, sonorance, manner, and four vowel features) derived exclusively from healthy control speech, and construct a 12-dimensional phonological profile. Evaluating 890 speakers across10corpora, 5 languages for the full MFA pipeline (English, Spanish, Dutch, Mandarin, French) and 3 primary aetiologies (Parkinsons disease, cerebral palsy, amyotrophic lateral sclerosis), we find that all five consonant d features correlate significantly with clinical severity (random-effects meta-analysis rho = -0.50 to -0.56, p < 2 x 10^-4; pooled Spearman rho = -0.47 to -0.55 with bootstrap 95% CIs not crossing zero), with the effect replicating within individual corpora, surviving FDR correction, and remaining robust to leave-one-corpus-out removal and alignment quality controls. Nasality d decreases monotonically from control to severe in 6 of 7 severity-graded corpora. Mann-Whitney U tests confirm that all 12 features distinguish controls from severely dysarthric speakers (p < 0.001).The method requires no dysarthric training data and applies to any language with an existing MFA acoustic model (currently 29 languages) or a model trained from healthy speech alone. It produces clinically interpretable per-feature profiles. We release the full pipeline and phone feature configurations for six languages to support replication and clinical adoption. Author SummaryOne of the authors has lived with ALS for sixteen years. Bernard Muller, who built this entire analytical pipeline using only eye-tracking technology, has experienced the progression of the disease firsthand, including the dysarthric speech that comes with advancing ALS and the tracheostomy that followed. The problem this paper addresses is not abstract to him, and that shapes how the method was designed. We developed a method to measure how well a person with dysarthria can produce distinct speech sounds, without needing any recordings of disordered speech for training. Our approach works by analysing how a widely available AI speech model organises different sound categories -- such as nasal versus oral consonants, or voiced versus voiceless sounds -- and measuring whether those categories become harder to tell apart. We tested this on 890 speakers across 10 datasets in five languages, covering Parkinsons disease, cerebral palsy, and ALS. Because the method only needs healthy speech recordings to set up, it applies to any language with an existing acoustic model, currently covering 29 languages. The resulting profiles show clinicians which specific aspects of speech production are degrading, rather than providing a single opaque severity score. This could support remote monitoring of speech decline in neurodegenerative disease and enable screening in languages and settings where specialist assessment is unavailable.

19
Planning Difficulties in Children and Adolescents with Hearing Loss across Development

Monteseirin, K.; Mendez-Couz, M.; Rivas-Fernandez, M. A.; Conejo, N. M.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.26.26361299 medRxiv
Top 0.1%
11.1%
Show abstract

Children and adolescents with hearing loss frequently encounter reduced auditory access and delayed language development, factors that may influence the maturation of executive functions. This study examined developmental differences in planning, a core executive function, in 98 children and adolescents with hearing loss or normal hearing aged 7 to18 years using the Tower of London task. Compared to normal hearing peers, participants with hearing loss made more unnecessary moves and rule violations and initiated problem-solving more rapidly, suggesting reduced preplanning efficiency and increased impulsivity. These group differences were most pronounced in adolescents, who showed faster initiation and greater movement inefficiency than age-matched normal hearing participants. Within the hearing loss group, adolescents displayed higher accuracy and longer initiation times than children, reflecting developmental improvements despite persistent gaps relative to hearing peers. Language development age did not alter the main effects. Findings indicate that reduced early auditory and language access may contribute to differences in planning development, highlighting the need for targeted executive functions support in educational and clinical settings for youth with hearing loss.

20
Towards an extended classification of noise-distortion preferences by modeling longitudinal dynamics of listening choices

Angonese, G.; Buhl, M.; Goesswein, J. A.; Kollmeier, B.; Hildebrandt, A.

2024-10-27 psychiatry and clinical psychology 10.1101/2024.10.25.24316092 medRxiv
Top 0.1%
10.9%
Show abstract

Individuals have different preferences for setting hearing aid (HA) algorithms that reduce ambient noise but introduce signal distortions. "Noise haters" prefer greater noise reduction, even at the expense of signal quality. "Distortion haters" accept higher noise levels to avoid signal distortion. These preferences were assumed to be stable over time, and individuals were classified solely on the basis of these stable, trait scores. However, the question remains as to how stable individual listening preferences are and whether day-to-day state-related variability needs to be considered as a further criterion for classification. We designed a mobile task to measure noise-distortion preferences over two weeks in an ecological momentary assessment study with N = 185 (106 f, Mage = 63.1, SDage = 6.5) unaided individuals with subjective hearing difficulties. Latent State-Trait Autoregressive (LST-AR) modeling was used to evaluate stability and dynamics of individual listening preferences. The analysis revealed a significant amount of state-related variance. The model has been extended to a mixture LST-AR model for data-driven classification, taking into account trait and state components of listening preferences. In addition to successful identification of noise haters, distortion haters and a third intermediate class based on longitudinal, outside of the lab data, we further differentiated individuals with different degrees of variability in listening preferences. It follows that individualisation of HA fitting could be improved by assessing individual preferences along the noise-distortion trade-off, and the day-to-day variability of these preferences needs to be taken into account for some individuals more than others.