Comparing AI- versus clinician-authored summaries of simulated primary care electronic health records
Shemtob, L.; Nouri, A.; Sullivan, A.; Qiu, C. S.; Martin, J.; Martin, M.; Noden, S.; Rob, T.; Neves, A. L.; Majeed, A.; Clarke, J.; Beaney, T.
Show abstract
Structured abstractO_ST_ABSObjectiveC_ST_ABSTo compare clinical summaries generated from simulated patient primary care electronic health records (EHRs) by ChatGPT-4 to summaries generated by clinicians on multiple domains of quality including utility, concision, accuracy and bias. Materials and MethodsSeven primary care physicians generated 70 simulated patient EHR notes, each representing 10 appointments with the practice over at least two years. Each record was summarised by a different clinician and by ChatGPT-4. AI- and clinician-authored summaries were rated blind by clinicians according to eight domains of quality and an overall rating. ResultsThe median time taken for a clinician to read through and assimilate the information in the EHRs before summarising, was seven minutes. Clinicians rated clinician-authored summaries higher than AI-authored summaries overall (7.39 versus 7.00 out of 10; p=0.02), but with greater variability in clinician-authored summary ratings. AI and clinician-authored summaries had similar accuracy and AI-authored summaries were less likely to omit important information and more likely to use patient-friendly language. DiscussionAlthough AI-authored summaries were rated slightly lower overall compared with clinician-authored summaries, they demonstrated similar accuracy and greater consistency. This demonstrates potential applications for generating summaries in primary care, particularly in the context of the substantial time taken for clinicians to undertake this work. ConclusionThe results suggest the feasibility, utility and acceptability of using AI-authored summaries to integrate into EHRs to support clinicians in primary care. AI summarisation tools have the potential to improve healthcare productivity, including by enabling clinicians to spend more time on direct patient care.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 95%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 94%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 95%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 94%
- Essential Indicators of Quality in Primary Care Settings: An Evidence-Based, Structured, Expert Approach 94%
- The challenges of replication: a worked example of methods reproducibility using routinely collected healthcare data 93%
Similar papers in this journal
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 96%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.