Is it time for the neurologist to use Large Language Models in everyday practice?
Maiorana, N. V.; Marceglia, S.; Treddenti, M.; Tosi, M.; Guidetti, M.; Creta, M. F.; Bocci, T.; Oliveri, S.; Martinelli Boneschi, F.; Priori, A.
Show abstract
Large Language Models (LLMs) such as ChatGPT and Gemini are gaining momentum in healthcare for their diagnostic potential. However, their real-world applicability in specialized medical fields like neurology remains inadequately explored. The possibility to use these tools in everyday diagnostic practice relies on the evaluation of their ability to serve as support for the clinician in assessing the patient, understanding the possible diagnosis and design the diagnostic pathway. To this end, in this study we (1) examined the available literature on the evaluation of LLMs in neurology diagnosis in order to understand whether the methodologies applied were adequate to translate the use of LLMs in everyday practice, and (2) designed and performed an experiment to evaluate the diagnostic accuracy and clinical recommendations of ChatGPT-3.5 and Gemini compared to neurologists using real-world clinical cases presented following the everyday diagnostic practice. In the vast literature of LLMs application in neurology, only 24 studies reported experiences using LLMs in clinical neurology. The experiments reported showed a heterogeneous scenario of prompt engineering and input formats. At present, while responses using structured prompts were well documented, there is a lack of studies using real-world clinical scenarios, and everyday workflows and practice. We therefore conducted a real-world experiment using a cohort of 28 anonymized patient records from the neurology department of the ASST Santi Paolo e Carlo Hospital (Milan, Italy). Cases were presented to ChatGPT-3.5 and Gemini replicating the typical clinical workflows. Diagnostic accuracy and appropriateness of recommended diagnostic tests were assessed against discharge diagnoses and neurologists performance. Neurologists achieved a diagnostic accuracy of 75%, outperforming ChatGPT-3.5 (54%) and Gemini (46%). Both LLMs exhibited difficulties in nuanced clinical reasoning and over-prescribed diagnostic tests in 17-25% of cases. Despite their ability to generate structured recommendations, they struggled with complex or ambiguous presentations, requiring additional prompts in some cases. We can therefore conclude that LLMs have potential as supportive tools in neurology but they currently lack the depth required for nuanced clinical decision-making. The findings emphasize the need for further refinement of LLMs and the development of evaluation methodologies that reflect the complexities of real-world neurology practice.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
Similar papers in this journal
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 91%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 90%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
Similar papers in this journal
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
- Development of a knowledge translation platform for ataxia: Impact on readers and volunteer contributors 92%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
Similar papers in this journal
- Towards AI-based Precision Rehabilitation via Contextual Model-based Reinforcement Learning 90%
- Technology-aided assessment of functionally relevant sensorimotor impairments in arm and hand of post-stroke individuals 90%
- Multimodal immersive trail making - virtual reality paradigm to study cognitive-motor interactions 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.