Case-level artificial intelligence for multi-photo teledermatology submissions: development and internal validation using patient-submitted dermatology images
Patel, V. P.; Sheth, N.; Patel, A.; Patel, Y.
Show abstract
Background: Store-and-forward teledermatology commonly relies on several patient-submitted photographs of the same concern, but most dermatology artificial intelligence models classify single images independently. Objective: To develop and internally validate a case-level diagnostic-support model that aggregates multiple patient-submitted photographs for common dermatologic conditions. Methods: We conducted a retrospective diagnostic-modeling study using the Skin Condition Image Network, a public dataset of deidentified self-taken dermatology images from US adults. We curated 2,336 cases comprising 5,041 images across 10 common inflammatory, allergic, and infectious conditions. Cases were split at the submission level into training, validation, and held-out test sets. Frozen general-purpose and dermatology-specific encoders were compared with image-level classifiers and a gated-attention multiple instance learning model that generated one case-level output from 1-3 images. Results: The strongest image-level baseline, dermatology-specific embeddings with random forest classification, achieved macro/micro ROC-AUCs of 0.797/0.854. Case-level aggregation improved discrimination, with dermatology-specific embeddings plus multiple instance learning achieving mean macro/micro ROC-AUCs of 0.819/0.863 across repeated stratified experiments. The locked final model achieved macro/micro ROC-AUCs of 0.800/0.849 on the held-out test set. Balanced-threshold sensitivity/specificity examples were 0.702/0.688 for eczema and 0.818/0.826 for urticaria. Limitations: Internal validation used a 10-condition subset from a US volunteer dataset; external validation, calibration, subgroup performance analysis, and prospective workflow studies are required. Conclusion: Modeling the teledermatology submission as a multi-image case better reflects asynchronous dermatology workflow than single-image classification. The model is preliminary clinician-facing support for structured review and triage, not autonomous diagnosis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 92%
- Zero Shot Health Trajectory Prediction Using Transformer 92%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 92%
Similar papers in this journal
Similar papers in this journal
- A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis 91%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 87%
- Point-of-care digital cytology with artificial intelligence for cervical cancer screening at a peripheral clinic in Kenya 87%
Similar papers in this journal
- ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images 92%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 92%
- Deep representation learning for clustering longitudinal survival data from electronic health records 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.