A 3-Minute Education on the False Positive Paradox Improves Trust Calibration in AI-Assisted Intracranial Aneurysm Detection: A Multinational Randomized Controlled Reader Study
Kim, S. H.; Le Guellec, B.; Rossmueller, P.; Schramm, S.; Boese, L.; Nikoubashman, O.; Kottlors, J.; Lichtenstein, T.; Strotzer, Q.; Meddeb, A.; Ziegelmeyer, S.; Steinhelfer, L.; Prucker, P.; Berberich, C.; Canisius, J.; Kreutzinger, V.; Hartl, F.; Schmitzer, L.; Rosenkranz, E.; Leonhardt, Y.; Beutel, T.-M.; Bitzer, F.; Maegerlein, C.; Boeckh-Behrens, T.; Baum, T.; Makowski, M. R.; Kirschke, J. S.; Bressem, K. K.; Adams, L. C.; Baird, G. L.; Wiestler, B.; Hedderich, D. M.
Show abstract
Background Even a highly accurate diagnostic test can yield more false-positive than true-positive findings in low-prevalence settings, which is known as the false positive paradox. Radiologists' unawareness of this paradox may foster automation bias, the tendency to excessively rely on AI outputs. Methods In this prospective, multinational, randomized controlled reader study (DRKS00038740), 34 readers from 10 countries (16 residents, 8 general radiologists or fellows, and 10 neuroradiologists) were randomly assigned to a control group (n = 17) or intervention group (n = 17), stratified by experience level. The intervention group reviewed a short, 3-minute educational video explaining the false positive paradox prior to the reading session. Both groups evaluated 20 TOF-MRA studies with AI-flagged findings (10% true-positive, 90% false-positive). Primary outcomes were acceptance rate of false-positive AI findings and follow-up intensity. These were evaluated using mixed models with crossed random effects for reader and case. Results At baseline, readers vastly overestimated the positive predictive value of AI tools for intracranial aneurysm detection (mean estimate, 62.9%; simulation-based estimate, 15.4% [95% interval, 8.1-28.0%]). The intervention reduced the odds of accepting AI false positives (OR 0.50 [upper 95% confidence bound, 0.95], one-sided p = 0.017), with acceptance probabilities of 12.7% (95% CI, 6.0-25.0%) in the intervention group compared to 22.5% (95% CI, 11.6-39.2%) in the control group. The intervention group exhibited a downward shift in follow-up intensity for false positives (OR 0.47 [upper 95% confidence bound, 0.81]; one-sided p = 0.014), recommending follow-up in 39.2% (120/306) of cases, compared to 54.9% (168/306) in the control group. Conclusion A brief education on the false positive paradox improved trust calibration in AI-assisted intracranial aneurysm detection. Our findings highlight the potential of reader-side cognitive debiasing strategies to improve trust calibration and support safer use of AI in radiology.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 94%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 92%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 89%
Similar papers in this journal
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 90%
- Development of a Novel Index to Characterise Arterial Dynamics Using Ultrasound Imaging 89%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 88%
Similar papers in this journal
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 93%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 92%
- Does contrast-enhancement improve visualisation of lenticulostriate arteries in cerebral small vessel disease using time-of-flight magnetic resonance angiography at 7 Tesla? 91%
Similar papers in this journal
- Evaluation of an artificial intelligence model for identification of intracranial hemorrhage subtypes on computed tomography of the head 91%
- A hybrid simulation-based workshop improves knowledge and confidence in the management of hemorrhagic conversion of stroke among interventional neurology trainees 89%
- Applying a Random Forest Approach in Predicting Health Status in Carotid Artery Stenosis Patients 30 Days Post-Stenting 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.