Back

Validation of an AI for Skin Diseases in Korea and Global Usage Statistics

Han, S. S.; Cho, S. I.; Chang, S. E.; Kim, S. H.; Na, J.-i.

2025-01-28 dermatology
10.1101/2025.01.27.25321219 medRxiv
Show abstract

BackgroundTo address the diversity of skin conditions and the low prevalence of skin cancers, we curated a large dataset and collected real-world data, to evaluate the real-world performance. MethodsWe evaluated a neural network model using the NIA dataset (70 diseases; 152,443 images from 92,004 cases from 9 hospitals). The NIA dataset was used to calculate the algorithms sensitivity for detecting skin cancer. Global usage statistics (1,690,849 queries) were analyzed to estimate the algorithms specificity for malignancy predictions, assuming all predictions were false positives. A global reader test (61,066 questionnaires) was performed using the SNU dataset (80 diseases; 240 images). ResultsFor malignancy diagnosis (NIA dataset), the sensitivity of the algorithm was 78.2% based on the three predicted differentials. In global usage statistics, among predictions for tumorous conditions, the malignancy prediction rates from the three differentials were 12.0% in Korea and 10.0% globally. The algorithms predictions for benign tumors were most prevalent in Asia (55.5%), Oceania (46.8%), and Europe (46.5%), while predictions for infectious diseases were more common in Africa (17.1%) In the reader test (SNU dataset), sensitivity / specificity of the algorithm (87.5% / 91.0%) outperformed those of participants (Korea = 56.9% / 87.2%, global = 55.5% / 84.3%). ConclusionFor skin cancer diagnosis in Korea, the sensitivity and specificity of the algorithm were estimated to be 78.2% and 88.0%. Further research on appropriate indications is required for screening use, taking overdiagnosis into consideration.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.