Back

Using explainable artificial intelligence to identify linguistic biomarkers of amyloid pathology in primary progressive aphasia

Robertson, C.; Rezaii, N.; Hochberg, D.; Quimby, M.; Worlff, P.; Dickerson, B. C.

2024-05-05 health informatics
10.1101/2024.05.02.24306657 medRxiv
Show abstract

IntroductionRecent success has been achieved in Alzheimers disease (AD) clinical trials targeting amyloid beta ({beta}), demonstrating a reduction in the rate of cognitive decline. However, testing methods for amyloid-{beta} positivity are currently costly or invasive, motivating the development of accessible screening approaches to steer patients toward appropriate diagnostic tests. Here, we employ a pre-trained language model (Distil-RoBERTa) to identify amyloid-{beta} positivity from a short, connected speech sample. We further use explainable AI (XAI) methods to extract interpretable linguistic features that can be employed in clinical practice. MethodsWe obtained language samples from 74 patients with primary progressive aphasia (PPA) across its three variants. Amyloid-{beta} positivity was established through the analysis of cerebrospinal fluid, amyloid PET, or autopsy. 51% of the sample was amyloid-positive. We trained Distil-RoBERTa for 16 epochs with a batch size of 6 and a learning rate of 5e-5, and used the LIME algorithm to train interpretation models to interpret the trained classifiers inference conditions. ResultsOver ten runs of 10-fold cross-validation, the classifier achieved a mean accuracy of 92%, SD = 0.01. Interpretation models were able to capture the classifiers behavior well, achieving an accuracy of 97% against classifier predictions, and uncovering several novel speech patterns that may characterize amyloid-{beta} positivity. DiscussionOur work improves previous research which indicates connected speech is a useful diagnostic input for prediction of the presence of amyloid-{beta} in patients with PPA. Further, we leverage XAI techniques to reveal novel linguistic features that can be tested in clinical practice in the appropriate subspecialty setting. Computational linguistic analysis of connected speech shows great promise as a novel assessment method in patients with AD and related disorders.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.