Skin Cancer Classification Using Explainable Artificial Intelligence With an Ensemble Model and Rigorous Leakage Free Validation
BARAN, M. T.; KARAKOYUN, O.
Show abstract
Background: Reliable melanoma classification requires models that capture both local dermoscopic morphology and broader contextual patterns while maintaining auditable, leakageaware internal validation. Objectives: To develop and internally validate an EfficientNetB0-Swin Transformer Tiny ensemble for classifying histopathologically verified dermoscopic images as benign melanocytic lesions or malignant melanoma. Methods: This retrospective diagnostic model-development and internal validation study screened 552,869 ISIC Archive records; filtering and dermatologist review yielded 1,199 uniquepatient and unique lesion images (578 benign and 621 malignant). Images were the predictors and histopathology was the reference. ImageNet pretrained EfficientNetB0 and Swin-T features were fused. Patient independent five fold validation used weighted sampling, mixup, label smoothing, AdamW, early stopping, and five view test time augmentation. Results: Mean accuracy was 0.89325 {+/-} 0.03179, mean receiver operating characteristic area under the curve (ROC-AUC) was 0.96348 {+/-} 0.01695, and mean support weighted F1-score was 0.89300 {+/-} 0.03220. The fold level 95% confidence intervals were 0.8538-0.9327 for accuracy and 0.9424-0.9845 for ROC-AUC. Pooled counts were 526 true negatives, 52 false positives, 76 false negatives, and 545 true positives, yielding 87.76% sensitivity and 91.00% specificity. Qualitative Grad-CAM review showed peripheral artifact activation in two false positives and lesion centered activation in two correctly classified cases; these observations were not systematically scored. Limitations: The validation folds were also used for early stopping and checkpoint selection. Device stratified analysis, systematic interpretability scoring, calibration, and independent external validation were unavailable. Conclusions: The ensemble showed high internal discrimination and is intended only as a clinician facing adjunct. The error audit workflow enables targeted retrospective review, but external validation is required before clinical use or generalizability claims.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards implementation of AI in New Zealand national screening program: Cloud-based, Robust, and Bespoke 91%
- Classification performance bias between training and test sets in a limited mammography dataset 91%
- Prediction of the ectasia screening index from raw Casia2 volume data for keratoconus identification by using convolutional neural networks 90%
Similar papers in this journal
Similar papers in this journal
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 92%
- PathProfiler: Automated Quality Assessment of Retrospective Histopathology Whole-Slide Image Cohorts by Artificial Intelligence, A Case Study for Prostate Cancer Research 92%
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 91%
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 92%
- Self-supervised contrastive learning improves machine learning discrimination of full thickness macular holes from epiretinal membranes in retinal OCT scans 91%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 90%
Similar papers in this journal
- Breast invasive ductal carcinoma classification on whole slide images with weakly-supervised and transfer learning 91%
- piNET: An Automated Proliferation Index Calculator Framework for Ki67 Breast Cancer Images 89%
- Accurate Screening for Early-Stage Breast Cancer by Detection and Profiling of Circulating Tumor Cells 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.