SHAPE AI: Development and Expert Validation of a Survey for Human AI Performance Evaluation in Healthcare
Mai, M. V.; Ozkaynak, M.; Marquard, J. L.; Simonson, R. J.; Holden, R. J.; Barton, H. J.; Catchpole, K.; Jamil, M.; Brauer, S.; Kandaswamy, S.
Show abstract
ObjectiveTo develop and content-validate a brief, expert-informed Survey for Human-AI Performance Evaluation (SHAPE-AI) for near-real-time assessment of how clinical AI affects human performance. BackgroundAI-enabled clinical decision support can improve outcomes only when aligned with clinician workflows, and cognitive demands. Existing evaluations measure technical performance and adoption, providing limited assessment of how AI shapes human performance. There is a lack of concise, operationally feasible instruments to measure AI impact on these human factors outcomes in clinical settings. MethodWe used a construct-driven, multi-stage development process. A literature review and prior qualitative work with users of a deployed sepsis prediction tool identified core human performance constructs. Preliminary items were created and iteratively refined through two expert panels. Six clinical informatics experts evaluated representativeness and clarity using content validity indices (CVI). Seven human factors experts then refined constructs, item wording, and response formats through ratings and focus groups, emphasizing discriminant validity, cognitive bias mitigation, and feasibility for deployment within 24 hours of AI use. ResultsA concise 10-item instrument was created, comprising perceived impact, interpretability, agreement with the AIs findings, agreement with the AIs recommendations, trust, workload, provider-patient and provider-team relationships, unexpected outcomes. ConclusionThe SHAPE-AI instrument is a theoretically grounded, operationally feasible tool to monitor human performance as relates to AI use. ApplicationHealth care organizations can deploy SHAPE-AI as a rapid, standardized probe to detect workflow misalignment, mis-calibrated reliance, communication disruptions, and unintended consequences of AI, informing safer design, implementation, and optimization of clinical AI tools. PrecisSHAPE-AI is a brief, expert-validated survey designed to capture clinicians near-real-time perceptions of how AI-enabled decision support affects human performance such as their situational awareness, decision-making, workload, and trust. SHAPE-AI offers health systems a practical, standardized way to monitor and understand the impact of AI on human performance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 95%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 95%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 93%
Similar papers in this journal
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 94%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 93%
Similar papers in this journal
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 94%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.