Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal
Shrestha, L.; Maharjan, D.; Bista, U.
Show abstract
Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design: Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the STARD 2015 checklist and STARD-AI/DECIDE-AI guidelines. Setting: Department of Radiology and Imaging, Patan Academy of Health Sciences / Patan Hospital, Lalitpur, Nepal. Participants: 826 consecutive health assessment applicants (foreign employment Pre-Departure Medical Examination and student migration) undergoing chest radiography between 5 June and 20 June 2026. Two cases were excluded due to DICOM technical failure. Index test: DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post-hoc derived threshold of 0.6258 (selected as the highest threshold achieving the pre-specified >=95% sensitivity criterion). Reference standard: Single-reader-per-case review by one of three radiologists - two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases) - each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings. Results: Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post-hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9-98.7%), specificity 77.2% (95% CI 74.1-80.0%), area under the receiver operating characteristic curve (AUROC) 0.9583 (95% bootstrap CI 0.9225-0.9843), NPV 99.67% (95% Wilson CI 98.8-99.9%), PPV 17.89% (95% Wilson CI 13.4-23.5%), and Cohen's {kappa} 0.237 (95% bootstrap CI 0.174-0.304). Brier score was 0.3621 (null Brier 0.0472) and ECE was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. Cross-validated results: Ten-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9-98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7-78.7%; optimism +1.40 pp), confirming primary metrics are not materially inflated by circular optimisation. Conclusions: The DenseNet-121 algorithm demonstrated high sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a rule-out tool (NPV 99.67%). Systematic score compression - preserved discrimination despite calibration shift - is a quantifiable marker of LMIC distributional shift. Prospective local calibration studies are warranted before operational deployment.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 97%
- A Comparison of CXR-CAD Software to Radiologists in Identifying COVID-19 in Individuals Evaluated for Sars CoV 2 Infection in Malawi and Zambia 96%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
Similar papers in this journal
- Early user experience and lessons learned using ultra-portable digital X-ray with computer-aided detection (DXR-CAD) products: A qualitative study from the perspective of healthcare providers 93%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 92%
Similar papers in this journal
- Point-of-care lung ultrasonography for early identification of mild COVID-19: a prospective cohort of outpatients in a Swiss screening center 93%
- Chest X-Ray Has Poor Diagnostic Accuracy and Prognostic Significance in COVID-19: A Propensity Matched Database Study 91%
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 90%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 94%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 94%
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 93%
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.