Deep learning-based computer-aided diagnostic models versus other methods for predicting malignancy risk in CT-detected pulmonary nodules
Watkins, J.; Akram, A.; Benemile, J. G.; Kathyrn, R.; Croce, F.; Wulaningsih, W.
Show abstract
ImportanceThere has been growing interest in the use of artificial intelligence (deep learning) to help achieve early diagnosis of prevalent diseases. None moreso than in lung cancer, where a combination of factors, including the high prevalence of nodules, the low prevalence of malignant nodules, and the indeterminacy of many nodules mean that it is fertile ground for the deployment of accurate, high-throughput deep learning (DL)-based tools. ObjectiveTo survey the landscape of externally validated DL-based computer-aided diagnostic (CADx) models, and assess their diagnostic performance for predicting the risk of malignancy in computed tomography (CT)-detected pulmonary nodules. Data sourcesAn electronic search was performed in the MEDLINE (PubMed), EMBASE, Science Citation Index, Cochrane Library databases (from inception to 10 April 2023). Study selectionStudies were deemed eligible if they were peer-reviewed experimental or observational articles that analysed the diagnostic performance of externally validated DL-based CADx models for the prediction of malignancy risk, with a direct comparison to models widely used in clinical practice. Data extraction and synthesisPRISMA guidelines were followed for the identification, screening, and selection process. A bivariate random-effect approach for the meta-analysis on the included studies was used. Quality Assessment of Diagnosis Accuracy Studies 2 (QUADAS-2) was used to assess risk of bias and applicability. Main outcomes and measuresMain outcomes included sensitivity, specificity, and area under the curve (AUC). ResultsAfter screening, 20 studies were included, comprising 7,664 participants and 10,128 nodules, of which 2,126 were malignant. DL-based CADx models were 15.8% more sensitive than physician judgement alone, and 35.4% more than clinical risk models alone. They had a similar pooled specificity as physician judgement alone (0.77 [95% CI: 0.69 -0.84] v 0.80 [95% CI: 0.71 -0.86], respectively), but were 5.5% more specific than clinical risk models alone. Accounting for threshold effects, DL-based CADx models had superior summary areas under the receiver operating characteristic curve (sAUROC), with relative sAUROCs of 1.06 (95% CI: 1.03-1.08) and 1.22 (95% CI: 1.19-1.24), as compared to physician judgement and clinical risk models alone, respectively. Conclusions and relevanceDL-based models show superior or comparable diagnostic performance when externally validated against widely used methods, such as the Brock and Mayo models. They have the potential to fulfil an unmet clinical-management need alongside experienced physician image readers. The included studies reported a high degree of heterogeneity, with threshold effects particularly prominent. Future research may consider more prospective studies and human-experimental studies. Key pointsO_ST_ABSQuestionC_ST_ABSHow effective are image-based, computer-aided diagnostic models that use deep learning methods to predict the malignancy risk of pulmonary nodules as compared with other methods used in clinical practice? FindingsThis systematic review and meta-analysis identified 20 observational studies (7,664 participants; 10,128 pulmonary nodules) from which pooled analyses found deep learning-based models to have a sensitivity of 0.88, specificity of 0.77, and summary area under the curve of 0.90 in predicting malignancy in pulmonary nodules. This was superior or comparable to other methods routinely used in clinical practice. MeaningDeep learning-based models are already being used in clinical practice in certain settings for nodule management. The results show their diagnostic performance justifies wider and more routine deployment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 94%
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 92%
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 89%
Similar papers in this journal
- Accuracy of deep learning based computed tomography diagnostic system of COVID-19: a consecutive sampling external validation cohort study 93%
- Classification performance bias between training and test sets in a limited mammography dataset 92%
- Quantitative analysis of chest computed tomography of COVID-19 pneumonia using a software widely used in Japan 92%
Similar papers in this journal
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 94%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 89%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 89%
Similar papers in this journal
- Chest X-Ray Has Poor Diagnostic Accuracy and Prognostic Significance in COVID-19: A Propensity Matched Database Study 92%
- Safety and feasibility of lung biopsy in diagnosis of acute respiratory distress syndrome: protocol for a systematic review and meta-analysis 92%
- Surgical Resection, Radiotherapy, And Percutaneous Thermal Ablation for Treatment of Stage 1 Non-Small Cell Lung Cancer: A Systematic Review and Network Meta-Analysis 92%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 90%
- Quantitative bias analysis methods for summary level epidemiologic data in the peer-reviewed literature: a systematic review 90%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.