Empathy and clarity in GPT-4-Generated Emergency Department Discharge Letters
Ben Haim, G.; Livne, A.; Manor, U.; hochstein, D.; Saban, M.; Blaier, o.; Abramov Iram, Y.; Gigi Balzam, M.; Lutenberg, A.; Eyade, R.; Qassem, R.; Trabelsi, D.; Dahari, Y.; Eisenmann, B. Z.; Shechtman, Y.; Nadkarni, G.; Glicksberg, B. S.; Zimlichman, E.; Perry, A.; klang, e.
Show abstract
Background and AimThe potential of large language models (LLMs) like GPT-4 to generate clear and empathetic medical documentation is becoming increasingly relevant. This study evaluates these constructs in discharge letters generated by GPT-4 compared to those written by emergency department (ED) physicians. MethodsIn this retrospective, blinded study, 72 discharge letters written by ED physicians were compared to GPT-4-generated versions, which were based on the physicians follow-up notes in the electronic medical record (EMR). Seventeen evaluators, 7 physicians, 5 nurses, and 5 patients, were asked to select their preferred letter (human or LLM) for each patient and rate empathy, clarity, and overall quality using a 5-point Likert scale (1 = Poor, 5 = Excellent). A secondary analysis by 3 ED attending physicians assessed the medical accuracy of both sets of letters. ResultsAcross the 72 comparisons, evaluators preferred GPT-4-generated letters in 1,009 out of 1,206 evaluations (83.7%). GPT-4 letters were rated significantly higher for empathy, clarity, and overall quality (p < 0.001). Additionally, GPT-4-generated letters demonstrated superior medical accuracy, with a median score of 5.0 compared to 4.0 for physician-written letters (p = 0.025). ConclusionGPT-4 shows strong potential in generating ED discharge letters that are empathetic and clear, preferable by healthcare professionals and patients, offering a promising tool to reduce the workload of ED physicians. However, further research is necessary to explore patient perceptions and best practices for leveraging the advantages of AI together with physicians in clinical practice.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A comparison of self-triage tools to nurse driven triage in the emergency department 96%
- Triaging and Referring In Adjacent General and Emergency Departments (the TRIAGE trial): a cluster randomised controlled trial 94%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 94%
Similar papers in this journal
- Variation in ambulance pre-alert process and practice: Cross-sectional survey of ambulance clinicians 94%
- Accuracy of the National Early Warning Score version 2 (NEWS2) in predicting need for time-critical treatment: Retrospective observational cohort study 94%
- The Balancing Act of Academic Clinical Fellows in UK Emergency Medicine: A Qualitative Study 92%
Similar papers in this journal
- A System Dynamics Model for Effects of Workplace Violence and Clinician Burnout on Agitation Management in the Emergency Department 92%
- Differences in medication reconciliation interventions between six hospitals: a mixed method study 92%
- Management of the COVID-19 health crisis: A survey in Swiss hospital pharmacies 91%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 92%
- Navigating the Crisis: A Cross-sectional Survey Analysis of Resident Doctors' Experiences of Specialty Training and Employment in Today's NHS 91%
- Changes in emergency department utilization in vulnerable populations after COVID-19 shelter in place orders 90%
Similar papers in this journal
- Understanding good communication in ambulance pre-alerts to Emergency Department. Findings from a qualitative study of UK emergency services 95%
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 93%
- Performance of the Safer Nursing Care Tool to measure nurse staffing requirements in acute hospitals: a multi-centre observational study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.