Back

Predicting Cognitive Function Using Transformer-Derived Speech Representations and Longitudinal Coherence Features

Devatha, D.; Xiao, J.

2026-08-06 neurology
10.64898/2026.08.04.26359730 medRxiv
Show abstract

Dementia affects more than 55 million people worldwide, and its progressive decline is difficult to track using infrequent in-person assessments, which can miss subtle changes between visits and add to clinician burden. One widely used measure, the Mini-Mental State Examination (MMSE), is administered intermittently and remains subject to inconsistent scoring judgment, limiting early detection. Prior speech-based machine learning approaches have largely focused on cross-sectional classification rather than longitudinal cognitive forecasting. We introduce a longitudinal, patient-level framework that combines transformer-derived semantic speech representations with longitudinal speech-change features and clinical history to forecast a patient's future MMSE score from their history of prior visits. To our knowledge, this is the first framework to unite transformer-derived speech encoding with longitudinal forecasting of cognitive severity, rather than single-visit classification alone. We evaluate this framework on longitudinal transcripts from the DementiaBank Pitt Corpus using a LightGBM gradient-boosted regression model, validated with patient-grouped cross-validation to prevent identity leakage between training and evaluation folds. The model forecasts future MMSE scores with high accuracy and stability across folds (R^2 = 0.840 +- 0.015, RMSE = 2.75, MAE = 2.02, Pearson r = 0.917), with transformer-derived speech representations contributing meaningful predictive signal alongside clinical history. Ablation analysis demonstrated that transformer-derived speech representations provided complementary predictive information beyond clinical variables. These results establish speech as a viable longitudinal digital biomarker of cognitive decline, offering a low-burden complement to intermittent clinical assessment that could enable earlier detection and more frequent monitoring, supporting better-timed care decisions for patients with dementia.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.