A 60-Second Interpretable Voice Model for Early Dementia Screening
Mekulu, K.; Aqlan, F.; Yang, H.
Show abstract
Early detection of cognitive impairment in assisted living is hindered by time-intensive tools like MMSE and MoCA. We present a 60-second voice-based screening model that analyzes picture descriptions to estimate dementia risk. Using transcripts from the DementiaBank corpus, our model integrates traditional linguistic features (pause rate, pronoun use, syntactic complexity) with latent semantic dimensions extracted from language model embeddings. These semantic axes, interpretable constructs like "Drift & Hesitation" or "Over-detailed Narration", consistently emerged as top predictors and may represent novel linguistic biomarkers of early decline. The final ElasticNet classifier is sparse, interpretable, and outperforms known non-deep learning baselines (AUC = 0.858), exceeding MMSE. Its simplicity enables deployment in mobile apps or in-room monitors, offering scalable, low-burden screening for early dementia. This work supports a shift toward linguistically grounded, tech-enabled cognitive care in aging populations. Author summaryEarly-stage dementia often goes undetected in assisted living communities, where time constraints and staffing limitations make routine cognitive screening impractical. While standard tools like the MMSE and MoCA require trained administration and take 10-15 minutes, our research introduces a fast, interpretable alternative: a 60-second voice-based screening model using picture description tasks. We analyzed speech samples from individuals describing a common scene, extracting both traditional linguistic features (e.g., pauses, pronouns, sentence complexity) and deeper thematic patterns using modern language embeddings. Our model not only outperforms widely used tools in accuracy, but also reveals interpretable language patterns such as hesitation or over-description that may serve as early signs of cognitive decline. Lightweight and explainable, this approach is well suited for mobile apps or in-room monitors, enabling scalable, low-burden dementia screening in real-world care environments.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CharMark: A Markov Approach to Linguistic Biomarkers in Dementia 95%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 92%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 88%
Similar papers in this journal
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 93%
- An aging focused unobtrusive and Privacy-Preserving Digital Behaviorome 92%
- A Remote Digital Memory Composite to Detect Cognitive Impairment in Memory Clinic Samples in Unsupervised Settings using Mobile Devices 91%
Similar papers in this journal
- Uncovering social states in healthy and clinical populations using digital phenotyping and Hidden Markov Models 92%
- Accurately Differentiating COVID-19, Other Viral Infection, and Healthy Individuals Using Multimodal Features via Late Fusion Learning 89%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 89%
Similar papers in this journal
- Evaluation of a speech-based AI system for early detection of Alzheimer’s disease remotely via smartphones 93%
- NeuropsychBrainAge: a biomarker for conversion from mild cognitive impairment to Alzheimer’s disease 92%
- Memorability of photographs in subjective cognitive decline and mild cognitive impairment: implications for cognitive assessment 92%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 91%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 90%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.