Back

Skin Cancer Classification Using Explainable Artificial Intelligence With an Ensemble Model and Rigorous Leakage Free Validation

BARAN, M. T.; KARAKOYUN, O.

2026-09-05 health systems and quality improvement
10.64898/2026.09.02.26362011 medRxiv
Show abstract

Background: Reliable melanoma classification requires models that capture both local dermoscopic morphology and broader contextual patterns while maintaining auditable, leakageaware internal validation. Objectives: To develop and internally validate an EfficientNetB0-Swin Transformer Tiny ensemble for classifying histopathologically verified dermoscopic images as benign melanocytic lesions or malignant melanoma. Methods: This retrospective diagnostic model-development and internal validation study screened 552,869 ISIC Archive records; filtering and dermatologist review yielded 1,199 uniquepatient and unique lesion images (578 benign and 621 malignant). Images were the predictors and histopathology was the reference. ImageNet pretrained EfficientNetB0 and Swin-T features were fused. Patient independent five fold validation used weighted sampling, mixup, label smoothing, AdamW, early stopping, and five view test time augmentation. Results: Mean accuracy was 0.89325 {+/-} 0.03179, mean receiver operating characteristic area under the curve (ROC-AUC) was 0.96348 {+/-} 0.01695, and mean support weighted F1-score was 0.89300 {+/-} 0.03220. The fold level 95% confidence intervals were 0.8538-0.9327 for accuracy and 0.9424-0.9845 for ROC-AUC. Pooled counts were 526 true negatives, 52 false positives, 76 false negatives, and 545 true positives, yielding 87.76% sensitivity and 91.00% specificity. Qualitative Grad-CAM review showed peripheral artifact activation in two false positives and lesion centered activation in two correctly classified cases; these observations were not systematically scored. Limitations: The validation folds were also used for early stopping and checkpoint selection. Device stratified analysis, systematic interpretability scoring, calibration, and independent external validation were unavailable. Conclusions: The ensemble showed high internal discrimination and is intended only as a clinician facing adjunct. The error audit workflow enables targeted retrospective review, but external validation is required before clinical use or generalizability claims.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.