Automated Detection of Early-Stage Dementia Using Large Language Models: A Comparative Study on Narrative Speech
Mekulu, K.; Aqlan, F.; Yang, H.
Show abstract
The growing global burden of dementia underscores the urgent need for scalable, objective screening tools. While traditional diagnostic methods rely on subjective assessments, advances in natural language processing offer promising alternatives. In this study, we compare two classes of language models--encoder-based pretrained language models (PLMs) and autoregressive large language models (LLMs) for detecting cognitive impairment from narrative speech. Using the DementiaBank Pitt Corpus and the widely used Cookie Theft picture description task, we evaluate BERT as a representative PLM alongside GPT-2, GPT-3.5 Turbo, GPT-4, and LLaMA-2 as LLMs. Although all models are pretrained, we distinguish PLMs and LLMs based on their architectural differences and training paradigms. Our findings reveal that BERT outperforms all other models, achieving 86% sensitivity and 95% specificity. LLaMA-2 follows closely, while GPT-4 and GPT-3.5 underperform in this structured classification task. Interestingly, LLMs demonstrate complementary strengths in capturing narrative richness and subtler linguistic features. These results suggest that hybrid modeling approaches may offer enhanced performance and interpretability. Our study highlights the potential of language models as digital biomarkers and lays the groundwork for scalable, AIpowered tools to support early dementia screening in clinical practice.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CharMark: A Markov Approach to Linguistic Biomarkers in Dementia 94%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 92%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 90%
Similar papers in this journal
- Practical Strategies for Extreme Missing Data Imputation in Dementia Diagnosis 94%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 92%
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 90%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 92%
- Using Artificial Intelligence to Learn Optimal Regimen Plan for Alzheimer’s Disease 92%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 91%
Similar papers in this journal
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Enhancing Fairness in Disease Prediction by Optimizing Multiple Domain Adversarial Networks 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.