Underdiagnosis Bias of Chest Radiograph Diagnostic AI can be Decomposed and Mitigated via Dataset Bias Attributions
Kobayashi, Y.; Joshi, S.; Zhang, H.; Singh, H.; Gichoya, J. W.
Show abstract
Inequitable diagnostic accuracy is a broad concern in AI-based models. However, current characterizations of bias are narrow, and fail to account for systematic bias in upstream data-collection, thereby conflating observed inequities in AI performance with biases due to distributional differences in the dataset itself. This gap has broad implications, resulting in ineffective bias-mitigation strategies. We introduce a novel retrospective model evaluation procedure that identifies and characterizes the contribution of distributional differences across protected groups that explain population-level diagnostic disparities. Across three large-scale chest radiography datasets, we consistently find that distributional differences in age and confounding image attributes (such as pathology type and size) contribute to poorer model performance across racial subgroups. By systematically attributing observed underdiagnosis bias to distributional differences due to biases in the data-acquisition process, or dataset biases, we present a general approach to disentangling how different types of dataset biases interact and compound to create observable AI performance disparities. Our method is actionable to aid the design of targeted interventions that recalibrate foundation models to specific subpopulations, as opposed to methods that ignore systematic contributions of upstream data biases on inequitable AI performance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Generative AI Enables Medical Image Segmentation in Ultra Low-Data Regimes 96%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 94%
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
Similar papers in this journal
- DeepAction: A MATLAB toolbox for automated classification of animal behavior in video 95%
- Construction and optimization of multi-platform precision pathways for precision medicine 94%
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.