Back

Better performance of deep learning pulmonary nodule detection using chest radiography with reference to computed tomography: data quality is matter

Kim, J. Y.; Ryu, W.-S.; Kim, D.; Kim, E. Y.

2023-02-14 radiology and imaging
10.1101/2023.02.09.23285621 medRxiv
Show abstract

BackgroundLabeling error may restrict radiography-based deep learning algorithms in screening lung cancer using chest radiography. Physicians also need precise location information for small nodules. We hypothesized that a deep learning approach using chest radiography data with pixel-level labels referencing computed tomography enhances nodule detection and localization compared to a data with only image-level labels. MethodsNational Institute Health dataset, chest radiograph-based labeling dataset, and AI-HUB dataset, computed tomography-based labeling dataset were used. As a deep learning algorithm, we employed Densenet with Squeeze-and-Excitation blocks. We constructed four models to examine whether labeling based on chest computed tomography versus chest X-ray and pixel-level labeling versus image-level labeling improves the performance of deep learning in nodule detection. Using two external datasets, models were evaluated and compared. ResultsExternally validated, the model trained with AI-HUB data (area under curve [AUC] 0.88 and 0.78) outperformed the model trained with NIH (AUC 0.71 and 0.73). In external datasets, the model trained with pixel-level AI-HUB data performed the best (AUC 0.91 and 0.86). In terms of nodule localization, the model trained with AI-HUB data annotated at the pixel level demonstrated dice coefficient greater than 0.60 across all validation datasets, outperforming models trained with image-level annotation data, whose dice coefficient ranged from 0.36-0.58. ConclusionOur findings imply that precise labeled data are required for constructing robust and reliable deep learning nodule detection models on chest radiograph. In addition, it is anticipated that the deep learning model trained with pixel-level data will provide nodule location information.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.