Pixel-Based Skin Tone Estimation on Dermoscopy: A Dual-Rater MST Benchmark and Feasibility Study
Kumarasinghe, A.; Bui, V.; Ghanbarzadeh, R.
Show abstract
Skin-tone labels are absent from public dermoscopy benchmarks such as the International Skin Imaging Collaboration (ISIC), making it impossible to audit whether clinical AI performs equitably across skin tones. While several recent works estimate skin tone automatically from clinical photography and selfies, we ask whether this approach is feasible on dermoscopy, the primary imaging modality of these benchmarks. To answer this, we make three main contributions. First, we release MST-Derm, a dual-rater Monk Skin Tone (MST) annotation benchmark on 500 ISIC 2018 images. Raters were given an explicit unrateable option for crops where the skin surrounding the lesion was too occluded to label confidently. We find that 60% of images were marked unrateable, yielding a 193-image consensus subset (quadratic-weighted Cohen's Kappa = 0.82). Second, we conduct a systematic feasibility study of three pixel-based MST annotation pipelines spanning the principal families in prior work: palette matching in perceptual colour space, robust colour statistics, and projection to a 1D colorimetric scalar. All three pipelines produce ordinal signal above chance (95% confidence intervals on quadratic-weighted Kappa exclude zero). However, ISIC 2018's extreme light-skin bias leaves 82% of the evaluation set at MST 2, giving a constant "always predict MST 2" baseline an accuracy floor the methods cannot overcome. To separate algorithmic signal from dataset bias, we evaluate on a class-balanced subset. The best method reaches quadratic-weighted Kappa = 0.43 against the trivial baseline of Kappa = 0.00, confirming the signal is genuine. Third, we diagnose this performance ceiling. We trace the bottleneck to two causes: dermoscopy's specialised illumination physically compresses the colour range on which lighter skin tones differ, and ISIC's dataset skew makes standard absolute-accuracy metrics uninformative. We conclude that while pixel-based colour features carry real MST signal on dermoscopy, current performance is insufficient for autonomous annotation. We release the benchmark, annotation protocol, all prediction runs, and analysis code to facilitate the development of robust skin-tone estimators, a vital prerequisite for accurately auditing fairness and mitigating bias in dermatological machine learning.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing generalizability of an AI-based visual test for cervical cancer screening 93%
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
Similar papers in this journal
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 93%
- Detection and measurement of butterfly eyespot and spot patterns using convolutional neural networks 93%
- Deep Learning Classification of Lipid Droplets in Quantitative Phase Images 92%
Similar papers in this journal
Similar papers in this journal
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 91%
- ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images 90%
- Large-scale capture of hidden fluorescent labels for training generalizable markerless motion capture models 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.