Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders
Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.
Show abstract
This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Effects of face masks on acoustic analysis and speech perception: Implications for peri-pandemic protocols 95%
- Representations of fricatives in sub-cortical model responses: comparisons with human consonant perception 95%
- A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception 95%
Similar papers in this journal
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 94%
- The Effect on Speech-in-Noise Perception of Real Faces and Synthetic Faces Generated with either Deep Neural Networks or the Facial Action Coding System 94%
- Auditory tests for characterizing hearing deficits in listeners with various hearing abilities: The BEAR test battery 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.