Back

NIR-SC-UFES: A portable NIR spectral dataset to skin cancer

da Cunha, P. H. P.; P. Zanoni, M.; D. Santos, F.; Tavares Nascimento, I.; Rezende, I.; R. P. Canuto, T.; de Paula Vieira, L.; C. S. Santos, M.; Romao, W.; H. L. Frasson, P.; Krohling, R.; R. Filgueiras, P.

2024-11-29 dermatology
10.1101/2024.11.27.24317165 medRxiv
Show abstract

In recent years, significant progress has been made in computer-aided diagnostics (CAD) for skin lesions, primarily using images and metadata. However, these methods have limitations, particularly in revealing the molecular structure of lesions. NIR spectroscopy offers additional data, capturing information not visible to the naked eye, which can improve automated CAD for skin lesions. Skin cancer remains a major cause of death, making early diagnosis crucial. A key challenge in applying machine and deep learning (MDL) techniques to spectroscopy is the lack of publicly available datasets. Previously, no public dataset of portable NIR spectral data for skin lesions existed. In collaboration with the Programa de Assistencia Dermatologica (PAD) at UFES, we developed a new dataset, NIR-SC-UFES, for skin cancer diagnosis. This dataset includes portable NIR spectral data for six lesion skin, with 714 spectra captured between 900 and 1700 nm. The dataset is available at data.mendeley.com/datasets/j9773cyr3k/1. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/24317165v1_ufig1.gif" ALT="Figure 1"> View larger version (34K): org.highwire.dtl.DTLVardef@1c7f3e3org.highwire.dtl.DTLVardef@5cc63aorg.highwire.dtl.DTLVardef@da2096org.highwire.dtl.DTLVardef@9198ea_HPS_FORMAT_FIGEXP M_FIG C_FIG Specifications Table O_TBL View this table: org.highwire.dtl.DTLVardef@ce3c61org.highwire.dtl.DTLVardef@1de06bdorg.highwire.dtl.DTLVardef@18ca3e0org.highwire.dtl.DTLVardef@5aecd9org.highwire.dtl.DTLVardef@173b35c_HPS_FORMAT_FIGEXP M_TBL C_TBL

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.