ChatGPT-o1 and the Pitfalls of Familiar Reasoning in Medical Ethics
Soffer, S.; Sorin, V.; Nadkarni, G.; Klang, E.
Show abstract
Large language models (LLMs) like ChatGPT often exhibit Type 1 thinking--fast, intuitive reasoning that relies on familiar patterns--which can be dangerously simplistic in complex medical or ethical scenarios requiring more deliberate analysis. In our recent explorations, we observed that LLMs frequently default to well-known answers, failing to recognize nuances or twists in presented situations. For instance, when faced with modified versions of the classic "Surgeons Dilemma" or medical ethics cases where typical dilemmas were resolved, LLMs still reverted to standard responses, overlooking critical details. Even models designed for enhanced analytical reasoning, such as ChatGPT-o1, did not consistently overcome these limitations. This suggests that despite advancements toward fostering Type 2 thinking, LLMs remain heavily influenced by familiar patterns ingrained during training. As LLMs are increasingly integrated into clinical practice, it is crucial to acknowledge and address these shortcomings to ensure reliable and contextually appropriate AI assistance in medical decision-making.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach? 94%
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 90%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 90%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
Similar papers in this journal
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 89%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 88%
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.