Effects of temperature settings on information quality of ChatGPT-3.5 responses: A prospective, single-blind, observational cohort study
Akamine, A.; Hayashi, D.; Tomizawa, A.; Nagasaki, Y.; Akamine, C.; Fukawa, T.; Hirosawa, I.; Saigo, O.; Hayashi, M.; Nanaoya, M.; Odate, Y.
Show abstract
ObjectiveThe effect of temperature settings on the quality of ChatGPT version 3.5 (OpenAI) responses related to drug information remains unclear. We investigated ChatGPT-3.5s response quality on apixaban information with and without the temperature being set to 0. MethodsOn 6 September 2023, 37 questions regarding apixaban, derived from the frequently asked questions on the Bristol-Myers Squibbs website, were entered into ChatGPT in Japanese. The primary endpoint was the effect of temperature settings on ChatGPT-3.5s responses to apixaban-related questions. The response accuracy, clarity, detail, and adequacy were rated on a 5-point Likert scale by 10 pharmacists, with higher scores indicating higher response quality. Cumulative score means were analyzed using the Mann-Whitney U test. In the subgroup analysis, evaluators were limited to pharmacists at university hospitals. Welchs t-test was employed in sensitivity analysis to validate primary endpoint findings. ResultsThe mean scores for ChatGPT-3.5s apixaban-related responses with (13.08) and without (14.40) the temperature being set to 0 were not significantly different (p = 0.064). Accuracy differed significantly (3.15 vs. 3.54, p = 0.045), whereas clarity, detail, and appropriateness were similar. Subgroup analysis (13.30 vs. 14.21, p = 0.394) and sensitivity analysis confirmed similar results (13.45 vs. 14.52, p = 0.105). ConclusionsChatGPT-3.5 temperature setting does not significantly affect overall responses to apixaban-related inquiries. However, the variance in accuracy suggests that ChatGPT-3.5 is unable to consistently provide precise responses. Hence, it is more suitable as a supplementary tool rather than a primary medical resource.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 92%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 92%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 93%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 92%
Similar papers in this journal
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 92%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 92%
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 92%
Similar papers in this journal
- Assessment of knowledge and perception of prescribers towards rational medicine use in the Ashanti Region of Ghana 94%
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 94%
- Perspectives of pharmacy employees on an inappropriate use of antimicrobials in Kathmandu, Nepal 93%
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Impact of electronic medical records on healthcare delivery in Nigeria: A Review 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.