Exploring Temperature Effects on Large Language Models Across Various Clinical Tasks
Patel, D.; Timsina, P.; Raut, G.; Freeman, R.; Levin, M.; Nadkarni, G.; Glicksberg, B. S.; Klang, E.
Show abstract
Large Language Models (LLMs) are becoming integral to healthcare analytics. However, the influence of the temperature hyperparameter, which controls output randomness, remains poorly understood in clinical tasks. This study evaluates the effects of different temperature settings across various clinical tasks. We conducted a retrospective cohort study using electronic health records from the Mount Sinai Health System, collecting a random sample of 1283 patients from January to December 2023. Three LLMs (GPT-4, GPT-3.5, and Llama-3-70b) were tested at five temperature settings (0.2, 0.4, 0.6, 0.8, 1.0) for their ability to predict in-hospital mortality (binary classification), length of stay (regression), and the accuracy of medical coding (clinical reasoning). For mortality prediction, all models accuracies were generally stable across different temperatures. Llama-3 showed the highest accuracy, around 90%, followed by GPT-4 (80-83%) and GPT-3.5 (74-76%). Regression analysis for predicting the length of stay showed that all models performed consistently across different temperatures. In the medical coding task, performance was also stable across temperatures, with GPT-4 achieving the highest accuracy at 17% for complete code accuracy. Our study demonstrates that LLMs maintain consistent accuracy across different temperature settings for varied clinical tasks, challenging the assumption that lower temperatures are necessary for clinical reasoning.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 96%
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 95%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 94%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 94%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.