Back

Evaluation of a Deep Learning and XAI based Facial Phenotyping Tool for Genetic Syndromes: A Clinical User Study

Sumer, O.; Huber, T.; Duong, D.; Ledgister Hanchard, S. E.; Conati, C.; Andre, E.; Solomon, B. D.; Waikel, R. L.

2025-06-09 genetic and genomic medicine
10.1101/2025.06.08.25328588 medRxiv
Show abstract

ObjectiveTo assess whether saliency-based explainable artificial intelligence (XAI) can improve diagnostic accuracy, confidence, and perceived usefulness for recognition of genetic conditions based on facial images. Materials and methodsForty-four medical geneticists, divided into AI-only (n=23) and XAI-supported (n=21) groups, assessed eighteen images of individuals with and without genetic conditions. Geneticists diagnostic accuracy and confidence were recorded before and after viewing an AI classifiers prediction probability, with or without XAI explanations. Mediation analyses were conducted to interpret geneticist behavior. ResultsAI-only and XAI-support improved geneticist accuracy for correct AI classifications, with mean improvements of 0.20{+/-}0.13 (AI-only) and 0.19{+/-}0.13 (XAI). Incorrect AI classifications decreased accuracy in both groups: -0.20{+/-}0.22 (AI-only) and -0.21{+/-}0.23 (XAI). Average confidence increased with correct classification: 0.30{+/-}0.26 (AI-only) and 0.38{+/-}0.71 (XAI) and decreased with incorrect classification: -0.14{+/-}0.64 (AI-only) and - 0.14{+/-}0.76 (XAI). AI prediction probability were generally considered useful (0.58{+/-}1.1 (AI-only) and 0.72{+/-}1.3 (XAI)); XAI explanations were viewed less favorably: -0.14{+/-}1.3 (saliency maps); -0.19{+/-}1.3 (region relevance scores). For incorrect AI classifications, there was a negative correlation between accuracy improvement and perceived AI usefulness (Spearmans Rho -0.31 (AI-only) and -0.40 (XAI)). Mediation analyses demonstrated that when AI is correct (without XAI), there was a significant mediated effect (0.583, 95% CI [0.044, 1.144]) by the model probability between user confidence and decision to follow AI. ConclusionsThe lack of accuracy or confidence improvements together with additional qualitative responses indicate that participants did not integrate saliency-based XAI into their decisions. AI prediction probability had a greater impact on participants decision making.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.