Prompting is All You Need: How to Make LLMs More Helpful for Clinical Decision Support
Dymm, B.; Goldenholz, D. M.
Show abstract
ImportanceLarge language models (LLMs) offer potential decision support, but their accuracy varies. Prompt engineering can generally enhance LLM behavior in a clinical context, yet best practices have yet to be formally explored in realistic neurology settings. ObjectiveTo evaluate the impact of structured prompting versus simple prompting on the performance of six LLMs (three closed-source: OpenAI GPT-4o, OpenAI o3, OpenAI GPT-5.2 Thinking; three open-source: Meta Llama-4-Scout-17B-16E-Instruct, Llama-3.3-70B-Instruct-Turbo, and the reasoning model R1-1776) for thrombolytic clinical decision support (CDS) in acute stroke. DesignModels responded to three novel ischemic stroke vignettes using either a simple question ("Should this patient be offered thrombolytics?") or a five-step structured prompt (CARDS) guiding information extraction, timing analysis, contraindication checking, decision process explanation, and risk-benefit discussion. Outputs were assessed across seven domains: guideline adherence, unsafe recommendations, risk recognition, guideline grading accuracy, inclusion of conversational explanation, clarity, and overall helpfulness. ResultsStructured prompts significantly enhanced performance across most domains, with varying effects between model families. For some closed-source models (GPT-4o, o3), prompts structured in the CARDS style improved guideline adherence from 83.3% to 100%, eliminated unsafe recommendations (16.7% to 0%), and increased specific guideline grading accuracy from 0% to 100%. The closed-source reasoning model GPT-5.2 Thinking similarly achieved 100% adherence, 0% unsafe recommendations, and 100% grading accuracy with structured prompts, while also maintaining perfect safety and risk recognition under simple prompting. Similarly, the open-source reasoning model R1-1776 achieved these top-tier outcomes (100% adherence, 0% unsafe, 100% grading, 100% conversation) when structured prompts were applied, with grading and conversation improving from 0%. In contrast, other open-source models (Llama-4-Scout, Llama-3.3-70B) showed more modest gains: risk recognition improved (83.3% to 100%) and guideline grading accuracy increased (0% to 66.7%), while guideline adherence (66.7%) and unsafe recommendations (33.3%) persisted. Overall, structured prompting yielded the largest improvements in guideline grading accuracy and conversational reasoning across multiple models. ConclusionStructured prompting substantially enhances LLM performance for acute stroke thrombolysis CDS. Notably, some models, including the proprietary GPT-4o, o3, and GPT-5.2 Thinking, and the open-source reasoning model R1-1776, achieved excellent safety and adherence with structured prompts. For clinical deployment of any LLM, structured prompts are crucial, and vigilant human oversight remains essential.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
Similar papers in this journal
- An Informatics Consult approach for generating clinical evidence for treatment decisions 92%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 92%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.