Geographic Domain Shift Precipitates Divergent Failure Modes In Deep Learning Based Tuberculosis Screening: A Multi-National External Validation Study
Shuaibu, I. I.; Khan, M. A.; Alkhamis, D.; Alkhamis, A.
Show abstract
BackgroundDeep learning algorithms for tuberculosis (TB) screening frequently achieve radiologist-level performance during internal evaluation, yet their reliability often degrades when deployed to populations differing from the training domain. Such degradation is clinically consequential for screening tools, where the World Health Organization (WHO) emphasizes high sensitivity to minimize missed infectious cases. MethodsA DenseNet-121 convolutional neural network was trained using transfer learning on the Shenzhen chest X-ray dataset (China; total n=662). To prevent anatomically implausible augmentation, horizontal flipping was excluded during training. The model was trained in two stages (head training followed by fine-tuning) and evaluated on: (i) an internal test set from China, (ii) an external balanced cohort from Montgomery County (USA; n=138), and (iii) an external TB-positive cohort from India (n=155). The India dataset served as a sensitivity stress test; specificity and ROC-AUC were not computed for this cohort due to the absence of negative controls. Model attention was explored using Grad-CAM. ResultsInternal validation yielded an Area Under the Curve (AUC) of 0.889 and accuracy of 85.6%. External testing revealed divergent failure modes. On the USA cohort, sensitivity was high (94.8%) but specificity decreased significantly (43.7%), indicating false-positive inflation. Conversely, on the India TB-only cohort, sensitivity collapsed to 52.3%, implying that 47.7% of confirmed TB cases were missed under domain shift. All metrics are reported as point estimates ConclusionGeographic domain shift produced non-uniform degradation false-positive surges in a low-burden setting and sensitivity collapse in a high-burden setting. These findings highlight the safety risks of deploying single-source TB screening AI without local validation and calibration.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 95%
- A Comparison of CXR-CAD Software to Radiologists in Identifying COVID-19 in Individuals Evaluated for Sars CoV 2 Infection in Malawi and Zambia 95%
- Classification of Hyper-scale Multimodal Imaging Datasets 95%
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 95%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 95%
- Automated Detection of COVID-19 through Convolutional Neural Network using Chest x-ray images 94%
Similar papers in this journal
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 95%
- COVID-Classifier: An automated machine learning model to assist in the diagnosis of COVID-19 infection in chest x-ray images 94%
- High-Dimensional Multinomial Multiclass Severity Scoring of COVID-19 Pneumonia Using CT Radiomics Features and Machine Learning Algorithms 93%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 94%
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- Point-of-care lung ultrasonography for early identification of mild COVID-19: a prospective cohort of outpatients in a Swiss screening center 91%
- Protocol for the development and validation of a machine-learning based tool for predicting the risk of hypertriglyceridemia in critically-ill patients receiving propofol sedation 91%
- Chest X-Ray Has Poor Diagnostic Accuracy and Prognostic Significance in COVID-19: A Propensity Matched Database Study 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.