Understanding Clinician Edits to Ambient AI Draft Notes: A Feasibility Analysis Using Large Language Models
Guo, Y.; Zhou, Y.; Hu, D.; Sutari, S.; Chow, E.; Tam, S.; Perret, D.; Pandita, D.; Zheng, K.
Show abstract
Ambient AI documentation tools generate draft notes that clinicians can review and edit before signing off in electronic health records. Scalable computational approaches to characterize how clinicians modify drafts remain limited, yet are essential for evaluating and improving AI effectiveness. We examined the feasibility of a few-shot prompted large language model (LLM) for categorizing sentence-level edits between AI drafts and final documentation. We developed five label-specific binary models targeting medication, symptom, diagnosis, orders/tests/procedures, and social history edits, and refined prompts using adversarial negatives and verification gates. Evaluation was performed against a human-annotated corpus. Medication and symptom models achieved promising performance (F1=0.787 and 0.780), whereas remaining models were precision-limited. Errors clustered in long, complex edits and category-boundary ambiguity. Therefore, prompt engineering is reliable for categorizing edits with explicit clues, while for complex context-dependent categories they are better suited for triage by labeling edits for human review.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 97%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 94%
Similar papers in this journal
Similar papers in this journal
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 93%
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 92%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 94%
- Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.