Temperature-Driven Variability in Emergency Diagnostic Accuracy by a Leading Language Model
Jarrett, P.; Hill, J.; Howell, M.; Grabow Moore, K.; Thoppil, J. J.; Vargas Ortiz, L.; Parnell, S. T.; Courtney, D. M.; McDonald, S. A.; Diercks, D. B.; Jamieson, A. R.; Cao, D.
Show abstract
ObjectiveTo determine the impact of the temperature parameter on GPT-4os diagnostic accuracy when evaluating emergency medicine cases and assess the effect on diagnostic divergence across iterations. MethodsWe conducted a simulation-based diagnostic accuracy study using four challenging emergency medicine cases adapted from the Foundations of Emergency Medicine curriculum. Each case was submitted to GPT-4o 250 times at five temperature settings (0.0, 0.25, 0.50, 0.75, 1.0), both with and without physical examination findings, yielding 10,000 total outputs. Each output contained exactly three differential diagnoses with one leading diagnosis. Diagnostic accuracy was assessed by comparing outputs against predetermined gold-standard diagnoses. ResultsAt temperature 0.0, GPT-4o achieved 100% leading diagnosis accuracy across all cases with physical exam data. As temperature increased, accuracy declined systematically to 89.4% at temperature 1.0. Diagnostic divergence increased dramatically from an average of 4.5 unique diagnoses at temperature 0.0 to 26.25 at temperature 1.0 (583% increase). Case sensitivity varied significantly, with ascending cholangitis showing the greatest temperature sensitivity (accuracy dropping from 100% to 70.4%) while carbon monoxide poisoning maintained 100% accuracy across all settings. DiscussionHigher temperatures introduced concerning diagnostic inconsistency rather than beneficial exploration, with substantial accuracy degradation in temperature-sensitive cases. ConclusionsLower temperature settings promote diagnostic accuracy and consistency, making them preferable for clinical applications requiring high reliability. Transparent reporting of temperature settings is essential for reproducible clinical artificial intelligence research. KEY MESSAGESO_ST_ABSWhat is already known on this topicC_ST_ABSLarge language models demonstrate promising diagnostic capabilities in medical reasoning tasks, but their non-deterministic nature and sensitivity to parameter settings remain poorly understood in clinical contexts. What this study addsThe temperature parameter significantly affects both diagnostic accuracy and consistency, with higher settings causing dramatic increases in diagnostic divergence. How this study might affect research, practice or policyThese findings mandate transparent reporting of temperature settings in clinical AI research and suggest that low-temperature configurations should be prioritized for high-reliability medical applications.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 92%
- ePOCT+ and the medAL-suite: Development of an electronic clinical decision support algorithm and digital platform for pediatric outpatients in low- and middle-income countries 92%
Similar papers in this journal
Similar papers in this journal
- Development and implementation of a clinical decision support system tool for the evaluation of suspected monkeypox infection 93%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
Similar papers in this journal
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.