Explainable AI as a Double-Edged Sword in Dermatology: The Impact on Clinicians versus The Public
Xu, X.; Hu, H.; Zhang, H.; Wang, W. K.; Wang, R.; Soenksen, L. R.; Badri, O.; Jafry, S.; Burger, E.; Nwandu, L.; Mehta, A.; Duhaime, E. P.; Qasim, A.; Lin, H.; Pereira, J.; Hershon, J.; Mui, P.; Gru, A. A.; Elhadad, N.; Mamykina, L.; Groh, M.; Tschandl, P.; Daneshjou, R.; Ghassemi, M.
Show abstract
Artificial intelligence (AI) is increasingly permeating healthcare, from physician assistants to consumer applications. Since AI algorithms opacity challenges human interaction, explainable AI (XAI) addresses this by providing AI decision-making insight, but evidence suggests XAI can paradoxically induce over-reliance or bias. We present results from two large-scale experiments (623 lay people; 153 primary care physicians, PCPs) combining a fairness-based diagnosis AI model and different XAI explanations to examine how XAI assistance, particularly multimodal large language models (LLMs), influences diagnostic performance. AI assistance balanced across skin tones improved accuracy and reduced diagnostic disparities. However, LLM explanations yielded divergent effects: lay users showed higher automation bias - accuracy boosted when AI was correct, reduced when AI erred - while experienced PCPs remained resilient, benefiting irrespective of AI accuracy. Presenting AI suggestions first also led to worse outcomes when the AI was incorrect for both groups. These findings highlight XAIs varying impact based on expertise and timing, underscoring LLMs as a "double-edged sword" in medical AI and informing future human-AI collaborative system design.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Understanding the robustness of vision-language models to medical image artefacts 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 92%
Similar papers in this journal
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 91%
Similar papers in this journal
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 93%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 92%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.