Development and initial evaluation of a conversational agent for Alzheimer's disease
Castano-Villegas, N.; Llano, I.; Villa, M. C.; Martinez, J.; Zea, J.; Urrea, T.; Banol, A. M.; Bohorquez, C.; Martinez, N.
Show abstract
BackgroundConversational Agents have attracted attention for personal and professional use. Their specialisation in the medical field is being explored. Conversational Agents (CA) have accomplished passing-level performance in medical school examinations and shown empathy when responding to patient questions. Alzheimers disease is characterized by the progression of cognitive and somatic decline. As the leading cause of dementia in the elderly, it is the subject of continuous investigations, which result in a constant stream of new information. Physicians are expected to keep up with the latest clinical guidelines; however, they arent always able to do so due to the large amount of information and their busy schedules. ObjectiveWe designed a conversational agent intended for general physicians as a tool for their everyday practice to offer validated responses to clinical queries associated with Alzheimers Disease based on the best available evidence. MethodologyThe conversational agent uses GPT-4o and has been instructed to respond based on 17 updated national and international clinical practice guidelines about Dementia and Alzheimers Disease. To approach the CAs performance and accuracy, it was tested using three validated knowledge scales. In terms of evaluating the content of each of the assistants answers, a human evaluation was conducted in which 7 people evaluated the clinical understanding, retrieval, clinical reasoning, completeness, and usefulness of the CAs output. ResultsThe agent obtained near-perfect performance in all three scales. It achieved a sensitivity of 100% for all three scales and a specificity of 75% in the less specific model. However, when modifying the input given to the assistant (prompting), specificity reached 100%, with a Cohens kappa of 1 in all tests. The human evaluation determined that the CAs output showed comprehension of the clinical question and completeness in its answers. However, reference retrieval and perceived helpfulness of the CA reply was not optimal. ConclusionsThis study demonstrates the potential of the agent and of specialised LLMs in the medical field as a tool for up-to-date clinical information, particularly when medical knowledge is becoming increasingly vast and ever-changing. Validations with health care experts and actual clinical use of the assistant by its target audience is an ongoing part of this project that will allow for more robust and applicable results, including evaluating potential harm.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Examining heterogeneity in dementia using data-driven unsupervised clustering of cognitive profiles 95%
- Top-funded digital health companies offering lifestyle interventions for dementia prevention: Company overview and evidence analysis 94%
- Random forest model for feature-based Alzheimer's disease conversion prediction from early mild cognitive impairment subjects 94%
Similar papers in this journal
- CharMark: A Markov Approach to Linguistic Biomarkers in Dementia 94%
- Medical Clinical Minds Meet Artificial Intelligence: Italian Physicians' Knowledge, Attitudes, and Concordance between Italian Physicians and AI-Generated Diagnoses. A National Cross-Sectional Study 91%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 90%
Similar papers in this journal
- Evaluating the GRACE voice assistant for dementia care among caregivers and healthcare professionals: An interview study. 95%
- User Experience Evaluation of Cogscreen for Screening Mild Cognitive Impairment: Formative and Summative Evaluation 94%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.