Back

Comparing AI- versus clinician-authored summaries of simulated primary care electronic health records

Shemtob, L.; Nouri, A.; Sullivan, A.; Qiu, C. S.; Martin, J.; Martin, M.; Noden, S.; Rob, T.; Neves, A. L.; Majeed, A.; Clarke, J.; Beaney, T.

2025-02-23 health informatics
10.1101/2025.02.21.25322674 medRxiv
Show abstract

Structured abstractO_ST_ABSObjectiveC_ST_ABSTo compare clinical summaries generated from simulated patient primary care electronic health records (EHRs) by ChatGPT-4 to summaries generated by clinicians on multiple domains of quality including utility, concision, accuracy and bias. Materials and MethodsSeven primary care physicians generated 70 simulated patient EHR notes, each representing 10 appointments with the practice over at least two years. Each record was summarised by a different clinician and by ChatGPT-4. AI- and clinician-authored summaries were rated blind by clinicians according to eight domains of quality and an overall rating. ResultsThe median time taken for a clinician to read through and assimilate the information in the EHRs before summarising, was seven minutes. Clinicians rated clinician-authored summaries higher than AI-authored summaries overall (7.39 versus 7.00 out of 10; p=0.02), but with greater variability in clinician-authored summary ratings. AI and clinician-authored summaries had similar accuracy and AI-authored summaries were less likely to omit important information and more likely to use patient-friendly language. DiscussionAlthough AI-authored summaries were rated slightly lower overall compared with clinician-authored summaries, they demonstrated similar accuracy and greater consistency. This demonstrates potential applications for generating summaries in primary care, particularly in the context of the substantial time taken for clinicians to undertake this work. ConclusionThe results suggest the feasibility, utility and acceptability of using AI-authored summaries to integrate into EHRs to support clinicians in primary care. AI summarisation tools have the potential to improve healthcare productivity, including by enabling clinicians to spend more time on direct patient care.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.