Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality
Takayama, T.; Sado, K.; Suda, K.; Tamura, H.; Ueda-Arakawa, N.; Ishihara, K.; Ueda, Y.; Okamoto, K.; Santos, L. H. d. O.; Oshika, T.; Kuroda, T.; Tsujikawa, A.; Miyake, M.
Show abstract
IMPORTANCELarge language models (LLMs) have been investigated for clinical documentation, with concerns about hallucinations and factual errors. Clinician review and revision of LLM-generated drafts are therefore considered essential, yet the impact of such workflow on both documentation time and quality remains unknown. OBJECTIVETo assess whether physician review and editing of LLM-generated drafts improves the time and quality of clinical documentation compared with clinician-only drafting in a randomized controlled trial setting. DESIGNSingle-center, parallel-group, prospective, randomized, open-label, blinded-endpoint (PROBE) pilot trial conducted from February 18 to March 14, 2025. SETTINGKyoto University Hospital, Department of Ophthalmology. PARTICIPANTSTwenty-one ophthalmology physicians were randomized; 17 completed the study, and 4 withdrew before initiating intervention. INTERVENTIONSParticipants were randomized to either the Clinician-in-the-loop group or the Clinician-only group. All participants created discharge summaries and referrals for six simulated patient records. In the Clinician-in-the-loop group, drafts were generated with an LLM assistant and then reviewed and edited by participants, whereas in the Clinician-only group, documents were drafted from scratch using matched templates. Unedited LLM drafts were additionally analyzed as the LLM-only group. MAIN OUTCOMES AND MEASURESDocument creation time (primary) and expert-rated document quality across six domains plus overall quality (secondary). RESULTSSeventeen physicians submitted 48 discharge summaries and 48 discharge referrals in the Clinician-in-the-loop group, 54 of each document type in the Clinician-only group, and 48 of each in the LLM-only group. For summaries, Clinician-in-the-loop was associated with shorter creation time versus Clinician-only ({beta} = -59.4 seconds; 95% CI, -118.1 to -0.8; P = .047). For referrals, clinician-in-the-loop required more time ({beta} = 94.8 seconds; 95% CI, 40.4 to 149.3; P<.001). In most quality domains for both document types, the clinician-in-the-loop workflow outperformed clinician-only drafting. LLM-only drafts were fastest but had the lowest quality. CONCLUSIONS AND RELEVANCEA clinician-in-the-loop approach improved document quality and accelerated documentation. Active clinician review of LLM-generated drafts is essential for clinical documentation, and such workflows may help enhance working conditions and patient care. TRIAL REGISTRATIONClinicalTrials.gov Identifier NCT07187050 Key PointsO_ST_ABSQuestionC_ST_ABSDoes a workflow in which clinicians review and edit LLM-generated drafts ("clinician-in-the-loop") improve the efficiency and quality of clinical documentation compared with clinician-only drafting? FindingsIn this single-center randomized controlled trial including 17 ophthalmology physicians, discharge summaries were significantly faster with the clinician-in-the-loop. Across both document types, clinician-in-the-loop drafts achieved higher quality scores than clinician-only documents. MeaningA clinician-in-the-loop workflow using LLMs can simultaneously enhance the efficiency and quality of clinical documentation; accordingly, active clinician review remains essential for improving working conditions and patient care overall.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
Similar papers in this journal
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.