Understanding Uncertainty in Large Language Model Predictions of Early Death in Critically Ill Patients: A Conformal Prediction Approach
Shah-mohammadi, F.; Millar, A.; Facelli, J.; Gouripeddi, R.
Show abstract
BackgroundEarly prediction of in-hospital death remains a significant challenge due to the limited availability of structured data during initial admission. Unstructured clinical notes, which often contain important observations and impressions, are an underutilized resource for real-time risk stratification. While leveraging recent advances in large language models (LLM) is a promising approach to use this unstructured information, the lack of understanding of the uncertainty of LLM predictions, at the patient level, for such critical forecasts is a serious deterrence for their use in clinical settings. ObjectiveThis study aims to evaluate the effectiveness and confidence, in predicting in-hospital death probability for an individual patient using LLMs, specifically GPT-4o and unstructured clinical notes. MethodsWe applied conformal prediction to quantify the uncertainty of GPT-4os zero-shot predictions for in-hospital death, leveraging concatenated clinical notes documented from the first 24 hours of intensive care unit (ICU) admission in MIMIC-III for patients with acute kidney failure who were admitted through the emergency department (ED). ResultsAcross both classes "in-hospital death" and "in-hospital survive", the GPT model performed better on the in-hospital death class, achieving precision 0.52 (95% CI 0.48-0.56), recall 0.93 (95% CI 0.90-0.95), and F1-score 0.66 (95% CI 0.63- 0.70). The conformal prediction (CP) framework provided an overall empirical coverage of 90.4%, exceeding the target threshold of 90%. However, class-specific coverage was imbalanced, with 99.7% for the death and 81.1% for the survived class. ConclusionsThe models outputs exhibit overconfidence, particularly in cases of incorrect predictions. Integrating conformal prediction provides a promising approach to quantifying and calibrating uncertainty in large language model outputs for individual patient predictions, thereby enhancing their potential applicability for clinical decision-making.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Creating a computer assisted ICD coding system: performance metric choice and use of the ICD hierarchy 97%
- Medication information extraction using local large language models 95%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 96%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 96%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.