Calibrating trust in AI-assisted pituitary surgery
Hudson, G. R.; Khan, D. Z.; Fayez, F.; Bhatia, S.; Bano, S.; Costanza, E.; Blandford, A.; Stoyanov, D.; McCulloch, P.; Marcus, H. J.; University College London Collaborators,
Show abstract
Background: Endoscopic endonasal transsphenoidal surgery (EETS) requires navigation around neurocritical anatomy. Today, artificial intelligence clinical decision support systems (AI-CDSSs) can orientate surgeons, but clinician trust in AI remains unclear, limiting safe deployment. This study evaluates how modifiable design affects trust and performance in a real-world pituitary surgery AI-CDSS. Method: Online, 70 clinicians with pituitary surgery experience were randomised evenly to a Basic or Enhanced AI-CDSS which outline the sella on EETS operative video. The Enhanced group additionally received explanation of the model and previous publications, alongside confidence labels depicting outline reliability. Both groups annotated the sella on six video clips, first alone then with the optional AI-CDSS. Clips were ordered by declining AI performance, except for the final clip. Self-reported trust was measured using a 1-7 scale after each annotation, and performance was the DICE overlap between user annotations and the ground truth. Comparisons used Mann-Whitney U and permutation analysis. Results: Sixty-four participants (91%) finished the exercise (31 Basic, 33 Enhanced). When AI performed best, median trust was 5.00 in both arms (U=559, p=.521). However, when AI performed worst, trust was significantly lower for the Enhanced group (3.00 vs 3.67, U=668, p=.035), sustained in the final clip (3.67 vs 4.33 U=687, p=.019). User performance improved with the AI-CDSS, but with no significant difference between the groups on the best or worst AI performing clips. Nevertheless, for the best AI, senior clinicians had higher median performance in the Enhanced group (0.95 vs 0.90, U=75, p=.066). There was also less dispersion in the Enhanced group when AI was inaccurate (IQR: 0.07 vs 0.21, p=.004). Conclusion: Interface design can improve trust calibration in a surgical AI-CDSS and may increment performance in seniors when AI is accurate, and consistency when AI is inaccurate. In future, these features may form important safety checks during translation to the operating room.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
- Artificial Intelligence for Surgical Scene Understanding: A Systematic Review and Reporting Quality Meta-Analysis 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
Similar papers in this journal
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 91%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 91%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
Similar papers in this journal
- Expert Surgeons and Deep Learning Models Can Predict the Outcome of Surgical Hemorrhage from One Minute of Video 93%
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 91%
- On evaluation metrics for medical applications of artificial intelligence 91%
Similar papers in this journal
- The roadmap for implementing value based healthcare in European university hospitals - consensus report and recommendations 87%
- Loon Lens 1.0 Validation: Agentic AI for Title and Abstract Screening in Systematic Literature Reviews 86%
- Statistical Decision Properties of Imprecise Trials Assessing COVID-19 Drugs 85%
Similar papers in this journal
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 92%
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 92%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.