Exploring the Interpretability of AI Decision Support Systems for Surgical Anatomy Recognition
Khan, D. Z.; Adams, T.; Wijekoon, A.; Ramirez Herrera, R.; Bano, S.; McCulloch, P.; Stoyanov, D.; Clarkson, M. J.; Costanza, E.; Blandford, A.; Marcus, H.; CARES Evaluation Group,
Show abstract
Artificial intelligence (AI) decision support systems for surgery hold promise but face barriers to adoption, particularly around the interpretability of their outputs. We conducted an international cross-sectional survey of 47 neurosurgeons to evaluate perspectives on literature-derived explanation techniques for AI-generated anatomical segmentations, using endoscopic pituitary surgery as a high-risk exemplar. Participants ranked certainty scores, certainty maps, saliency maps, scene similarity scores, and nearest-neighbour illustrations, and rated them using a modified Explanation Satisfaction Scale alongside free-text feedback. Certainty-based techniques were consistently ranked and rated highest for interpretability - valued for aligning with surgical decision-making by conveying confidence (via scores) and anatomical boundaries (via maps). Saliency- and similarity-based methods were judged less clinically relevant and better suited to educational settings. Certainty-based explanations, therefore, appear most acceptable to surgeons for clinical integration of decision support systems, though their impact on AI acceptability, trust calibration, and performance requires prospective evaluation across surgical domains.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 91%
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 91%
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 90%
Similar papers in this journal
- Expert Surgeons and Deep Learning Models Can Predict the Outcome of Surgical Hemorrhage from One Minute of Video 94%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 93%
- Development and Validation of Collaborative Robot-assisted Cutting Method for Iliac Crest Flap Raising: Randomized Crossover Trial 92%
Similar papers in this journal
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 93%
- Virtual reality as a strategy for intra-operatory anxiolysis and pharmacological sparing in patients undergoing breast surgeries: the V-RAPS randomized controlled trial protocol 92%
- The Impact of a Wireless Audio System on Communication in Robotic-Assisted Laparoscopic Surgery: A Prospective Controlled Trial 92%
Similar papers in this journal
- Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations 94%
- Performance of ChatGPT, GPT-4, and Google Bard on a Neurosurgery Oral Boards Preparation Question Bank 94%
- A Porcine Model of Peripheral Nerve Injury Enabling Ultra-Long Regenerative Distances: Surgical Approach, Recovery Kinetics, and Clinical Relevance 87%
Similar papers in this journal
- Surgery & COVID-19: A rapid scoping review of the impact of COVID-19 on surgical services during public health emergencies 92%
- Protocol of the observational study STRATUM-OS: First step in the development and validation of the STRATUM tool based on multimodal data processing to assist surgery in patients affected by intra-axial brain tumours 92%
- Validity of intraoperative imageless navigation (Naviswiss™) for component positioning accuracy in primary total hip arthroplasty: Protocol for a prospective observational cohort study in a single-surgeon practice 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.