An Experimental Investigation of the Relationship between AI-Human Workflow Design and Legal Liability for Radiologists: The Erroneous-Change Penalty and Omission Bias
Song, E. C.; Bernstein, M. H.; Sheppard, B.; Bruno, M. A.; Baird, G. L.
Show abstract
Background: With growing impetus to integrate artificial intelligence (AI) tools into radiology, clinical practices must navigate workflow redesign. This carries implications for medical malpractice liability. Methods: We conducted an online vignette experiment with United States adults who acted as hypothetical jurors in a malpractice case involving a missed intracranial hemorrhage. Participants (n=2,347) were randomized to one of 22 conditions: a no-AI control and 21 conditions involving a hypothetical AI system. These twenty-one conditions varied by whether (1) a single-read or double-read workflow was used, (2) the radiologist's initial interpretation was documented, (3) the radiologist changed their interpretation after viewing AI output, (4) the AI detected the abnormality, and (5) the AI error rate--False Discovery Rate (FDR) or False Omission Rate (FOR--was provided to participants only, both participants and radiologist, or neither. The primary outcome was perceived liability, assessed by whether the radiologist met their duty of care. Findings: Perceived liability differed across conditions (p<0.0001). Double-read workflows (p<0.0001), documenting initial interpretations (p=0.0125), and providing participants with AI error rates, including the FDR (p=0.0038) or FOR (p=0.0035), reduced perceived liability. Liability was also lower when AI was incorrect (p<0.0001). Radiologists' awareness of AI error rates did not significantly impact liability. Notably, we observed an erroneous change penalty: the greatest liability occurred when radiologists initially identified an abnormality but later changed their interpretation to normal after seeing that AI identified the case as normal; conversely, perceived liability was lowest with documented, double-read workflows. Interpretation: Double-read workflows with documented initial interpretations and disclosure of AI error rates reduce perceived liability, though changing a correct initial interpretation increases it. Strategic workflow design is critical for successful AI implementation that can mitigate malpractice risk.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 90%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 90%
- Classification performance bias between training and test sets in a limited mammography dataset 90%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 93%
- Expert Surgeons and Deep Learning Models Can Predict the Outcome of Surgical Hemorrhage from One Minute of Video 90%
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 90%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 91%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 90%
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- Simulated Misuse of Large Language Models and Clinical Credit Systems 91%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 90%
Similar papers in this journal
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 92%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 90%
- Crowdfunding Medical Care: A Comparison of Online Medical Fundraising in Canada, the United Kingdom, and the United States 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.