Back

A Machine Learning Classification of Individuals with Mild Cognitive Impairment into Variants from Writing

Themistocleous, C.; Kim, H.

2024-02-19 neurology
10.1101/2024.02.16.24302965 medRxiv
Show abstract

IntroductionIndividuals with Mild Cognitive Impairment (MCI), a transitional stage between cognitively healthy aging and dementia, are characterized by subtle neurocognitive changes. Clinically, they can be grouped into two main variants, namely into patients with amnestic MCI (aMCI) and non-amnestic MCI (naMCI). The distinction of the two variants is known to be clinically significant as they exhibit different progression rates to dementia. However, it has been particularly challenging to classify the two variants robustly. Recent research indicates that linguistic changes may manifest as one of the early indicators of pathology. Therefore, we focused on MCIs discourse-level writing samples in this study. We hypothesized that a written picture description task can provide information that can be used as an ecological, cost-effective classification system between the two variants. MethodsWe included one hundred sixty-nine individuals diagnosed with either aMCI or naMCI who received neurophysiological evaluations in addition to a short-written picture description task. Natural Language Processing (NLP) and BERT pre-trained Language Models were utilized to analyze the writing samples. ResultsWe showed that the written picture description task provided 90% overall classification accuracy for the best classification models, which performs better than cognitive measures. DiscussionWritten discourses analyzed the AI models can automatically assess individuals with aMCI and naMCI and facilitate diagnosis, prognosis, therapy planning, and evaluation.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.