Reproducible Generative AI Evaluation for Healthcare: A Clinician-in-the-Loop Approach
Livingston, L.; Featherstone-Uwague, A.; Barry, A.; Barretto, K.; Morey, T.; Herrmannova, D.; Avula, V.
Show abstract
ObjectiveTo develop and apply a reproducible methodology for evaluating generative artificial intelligence powered systems in healthcare, addressing the gap between theoretical evaluation frameworks and practical implementation guidance. Materials and MethodsA five dimension evaluation framework was developed to assess query comprehension and response helpfulness, correctness, completeness, and potential clinical harm. The framework was applied to evaluate ClinicalKey AI using queries drawn from user logs, a benchmark dataset, and subject matter expert curated queries. Forty one board certified physicians and pharmacists were recruited to independently evaluate query-response pairs. An agreement protocol using the mode and modified Delphi method resolved disagreements in evaluation scores. ResultsOf 633 queries, 614 (96.99%) produced evaluable responses, with subject matter experts completing evaluations of 426 query-response pairs. Results demonstrated high rates of response correctness (95.5%) and query comprehension (98.6%), with 94.4% of responses rated as helpful. Two responses (0.47%) received scores indicating potential clinical harm. Pairwise consensus occurred in 60.6% of evaluations, with remaining cases requiring third tie-breaker review. DiscussionThe framework demonstrated effectiveness in quantifying performance through comprehensive evaluation dimensions and structured scoring resolution methods. Key strengths included representative query sampling, standardized rating scales, and robust subject matter expert agreement protocols. Challenges emerged in managing subjective assessments of open-ended responses and achieving consensus on potential harm classification. ConclusionThis framework offers a reproducible methodology for evaluating healthcare generative artificial intelligence systems, establishing foundational processes that can inform future efforts while supporting the implementation of generative AI applications in clinical settings.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 96%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 97%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 95%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 94%
Similar papers in this journal
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 94%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 94%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.