Evaluating Accuracy and Reasoning Capabilities of Large Language Models for Acute Ischemic Stroke Management
Meddeb, A.; Bakhtiari, N.; Rangus, I.; Fetcher, L.; Le Guellec, B.; Busch, F.; Doucet, A.; Hua, V. T.; Verot-Nguyen, F.; Manceau, P.-F.; Paiusan, L. V.; PAGANO, P.; Wattjes, M. P.; Moulin, S.; Pierot, L.; Soize, S.
Show abstract
BackgroundAcute ischemic stroke (AIS) management has evolved substantially over the past two decades, with mechanical thrombectomy adding complexity that requires specialized centers. As many patients initially present to primary care facilities, rapihtmld and accurate triage is critical. Large language models (LLMs) may help bridge expertise gaps, especially where stroke specialists are not immediately available. This study evaluates the diagnostic accuracy and reasoning quality of four LLMs in determining eligibility for intravenous thrombolysis (IVT) and mechanical thrombectomy (MT), compared with experienced clinicians and real-world treatment decisions. MethodsWe retrospectively collected 80 acute ischemic stroke cases from two stroke centers. Cases were presented to LLMs as well to clinicians as clinical vignettes containing demographic, clinical, and imaging data. Four LLMs (DeepSeek R1, OpenAI o3 mini, Gemini 2.0, LLaMA 3.3) and six stroke experts (two neurologists, four neuroradiologists) independently reviewed the cases and recommended one or more treatment strategies including IVT and MT. The ground truth was defined as the institutional treatment decision. Accuracy for MT and IVT recommendations was calculated for both LLMs and clinicians. Additionally, a qualitative error analysis evaluated the reasoning ability of LLMs. ResultsOpen-source reasoning model DeepSeek R1 outperformed all other LLMs and clinicians for MT (87% accuracy) and achieved 78% accuracy for IVT. Across models, accuracy was generally higher for MT than for IVT. Neurologists reached 81% (MT) and 80% (IVT), while neuroradiologists achieved 84% (MT) and 76% (IVT). Reasoning analysis for MT recommendations showed that most errors were clinically reasonable but differed from real-world decision, whereas IVT errors were primarily due to guideline non-adherence. ConclusionsLLMs can match or even exceed expert clinician performance in MT and IVT eligibility decisions, while providing transparent reasoning. These findings support prospective evaluation of LLM-based decision support in acute stroke care, especially in settings without immediate specialist expertise.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 93%
- Machine learning-based forecasting of daily acute ischemic stroke admissions using weather data 91%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 90%
Similar papers in this journal
- Automated Identification of Thrombectomy Amenable Vessel Occlusion on Computed Tomography Angiography using Deep Learning 94%
- End to end stroke triage using cerebrovascular morphology and machine learning 93%
- Left Hemisphere Bias of NIH Stroke Scale is Most Severe for Middle Cerebral Artery Strokes 92%
Similar papers in this journal
- Large Core Thrombectomy: Feasibility Of Simplified Protocol In Resource-Limited Settings 94%
- A hybrid simulation-based workshop improves knowledge and confidence in the management of hemorrhagic conversion of stroke among interventional neurology trainees 93%
- Prehospital triage of intracranial hemorrhage and anterior large vessel occlusion ischemic stroke: the value of the rapid arterial occlusion evalution 93%
Similar papers in this journal
- Towards AI-based Precision Rehabilitation via Contextual Model-based Reinforcement Learning 90%
- Rest the Brain to Learn New Gait Patterns after Stroke 89%
- “It all ends too soon” - Exploring stroke survivors and physiotherapists perspectives on stroke rehabilitation and the role of technology for promoting access to rehabilitation in the community 88%
Similar papers in this journal
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 95%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
- Opening the Black Box of Artificial Intelligence for Clinical Decision Support: A Study Predicting Stroke Outcome 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.