Measuring Reliability in Locally-deployed Language Model Dysarthric Speech Assessments
Klempir, O.; Mullerova, J. G.; Tichopad, A.; Krupicka, R.
Show abstract
Speech is a rich and non-invasive source of clinical information, potentially providing digital biomarkers for neurological disorders such as Parkinsons disease (PD). Impaired articulation and reduced intelligibility are among the most pervasive PD symptoms, which has motivated research into automated, objective quantification of speech deficits. This study investigated whether metrics derived from automatic speech recognition (ASR) and large language models (LLMs) can quantify speech intelligibility and describe clinical severity. Recordings of fixed read text from patients with PD and healthy controls (HC) were transcribed and evaluated using conventional ASR error measures (such as Word and Character Error Rate), a proposed Mistral-based LLM intelligibility score, and BERT-derived typo metrics (BERT, short for Bidirectional Encoder Representations from Transformers). Group-level discriminability between PD (N = 16) and HC (N = 21) was low (Mann-Whitney p > 0.05; Random Forest obtained Receiver Operating Characteristic Area Under the Curve of 0.66; leave-one-subject-out evaluation), indicating that transcript-level features alone offer limited classification abilities. Importantly, the LLM-derived intelligibility score demonstrated excellent repeatability across five runs (intra-class correlation; ICC = 0.97, Cronbachs = 0.98) and showed strong correlations with ASR error metrics. Both LLM and ASR measures significantly correlated (p < 0.05; Spearmans rank correlation) with common rating clinical scales (Hoehn and Yahr scale and Unified Parkinsons Disease Rating Scale), whereas BERT-typo parameters did not. These findings suggest the use of LLMs as tools for generating reference-free intelligibility scores that reflect disease severity.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Remote monitoring of progression in early Parkinson’s disease: reliability and validity of the Roche PD Mobile Application v2 93%
- Developing and Validating a New Web-Based Tapping Test for Measuring Distal Bradykinesia in Parkinson's Disease 92%
- CDS-PD: A Novel Clinical Decision Support Platform for Parkinson's Disease 91%
Similar papers in this journal
Similar papers in this journal
- Motor signatures in digitized cognitive and memory tests enhances characterization of Parkinson’s disease 95%
- Analyzing wav2vec embedding in Parkinson’s disease speech: A study on cross-database classification and regression tasks 95%
- An explainable spatial-temporal graphical convolutional network to score freezing of gait in parkinsonian patients 93%
Similar papers in this journal
- Digital risk score sensitively identifies presence of α-synuclein aggregation or dopaminergic deficit 92%
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 90%
- α-Synuclein seeding activity and progression in sporadic and genetic forms of Parkinson's disease in the Parkinson's Progression Markers Initiative cohort 89%
Similar papers in this journal
- Remote Real Time Digital Monitoring fills a Critical Gap in the Management of Parkinson’s disease 93%
- Identification and prediction of Parkinson's disease subtypes and progression using machine learning in two cohorts. 92%
- Digital markers from smartwatch data relate to non-motor clinical examinations of Parkinson’s disease 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.