Pretraining Diversity and Clinical Metric Optimization Achieve State-of-the-Art Performance on ChestX-ray14
Fisher, G. R.
Show abstract
We achieved state-of-the-art performance on the NIH ChestX-ray14 multi-label classification task using a simple 3-model ensemble: mean ROC-AUC 0.940, F1 0.821 (95% CI: 0.799-0.845), PR-AUC 0.827, sensitivity 76.0%, and specificity 98.8% across 14 thoracic diseases. Our primary finding challenges current research priorities: pretraining diversity dominates architectural diversity. Systematic evaluation of 255 ensemble combinations from 8 models spanning three architecture families (ConvNeXt, Vision Transformers, EfficientNet) at multiple resolutions (224x224 to 384x384) revealed that a simple 3-model ConvNeXt ensemble combining ImageNet-1K, ImageNet-21K, and ImageNet-21K-384 pretrained variants outperformed all 252 alternative combinations, including modern Vision Transformers and efficiency-optimized architectures. This ensemble achieved mean ROC-AUC 0.940, exceeding recent hybrid transformer approaches (LongMaxViT [1]: 0.932) with substantially lower computational requirements. Systematic comparison of five optimization strategies (F1, F_SS, pure sensitivity, Youdens J, validation loss) established that clinical metric optimization outperforms traditional validation loss by 19.5% in F1 score. F_SS optimization (sensitivity-specificity harmonic mean) achieved optimal clinical balance: highest sensitivity (73.9%), best Youdens J (0.727), and superior threshold-independent performance (ROC-AUC, PR-AUC). Traditional validation loss optimization failed to align with diagnostic utility despite achieving mathematical convergence. Strategic pretraining selection and clinical metric optimization provide greater performance improvements than architectural innovation alone, enabling competitive state-of-the-art results on accessible computational resources (AWS g5.2xlarge, $1.21/hr).
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 97%
- Classification of Hyper-scale Multimodal Imaging Datasets 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
Similar papers in this journal
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 95%
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 95%
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 95%
Similar papers in this journal
- The Effect of Image Resolution on Automated Classification of Chest X-rays 97%
- Predicting Primary Site of Secondary Liver Cancer with a Neural Estimator of Metastatic Origin (NEMO) 95%
- A 3D CNN Classification Model for Accurate Diagnosis of Coronavirus Disease 2019 using Computed Tomography Images 93%
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 96%
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 96%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.