Quantifying Explainable AI Methods in Medical Diagnosis: A study in skin cancer
Sangwan, H.
Show abstract
Deep learning models have shown substantial promise in assisting medical diagnosis, offering the potential to improve patient outcomes and reduce clinician workloads. However, the widespread adoption of these models in clinical practice has been hindered by concerns surrounding their trustworthiness, transparency, and interpretability. Addressing these challenges requires not only the development of explainable AI (xAI) techniques but also quantitative metrics to evaluate their effectiveness. This study presents a comprehensive framework for training, explaining, and quantitatively assessing deep learning models for skin cancer diagnosis. Leveraging the HAM10000 dataset of seven diagnostic skin lesion categories, multiple convolutional neural network architectures--including custom CNNs, DenseNet, MobileNet, and ResNet--were trained and optimized using augmentation, oversampling, and hyperparameter tuning. Following model training, explainability techniques such as SHAP, LIME, and Integrated Gradients were deployed to generate post hoc explanations. Critically, the primary contribution of this work is the quantitative evaluation of these explanation methods using metrics related to faithfulness, robustness, and complexity. All code, models, and results are publicly available, providing a reproducible pathway toward more trustworthy, explainable diagnostic tools.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 96%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 95%
- The mathematics of erythema: Development of machine learning models for artificial intelligence assisted measurement and severity scoring of radiation induced dermatitis 94%
Similar papers in this journal
Similar papers in this journal
- Equipping Computational Pathology Systems with Artifact Processing Pipelines: A Showcase for Computation and Performance Trade-offs 94%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.