Discovering mechanisms underlying medical AI prediction of protected attributes
Gadgil, S. U.; DeGrave, A. J.; Daneshjou, R.; Lee, S.-I.
Show abstract
Recent advances in Artificial Intelligence (AI) have started disrupting the healthcare industry, especially medical imaging, and AI devices are increasingly being deployed into clinical practice. Such classifiers have previously demonstrated the ability to discern a range of protected demographic attributes (like race, age, sex) from medical images with unexpectedly high performance, a sensitive task which is difficult even for trained physicians. In this study, we motivate and introduce a general explainable AI (XAI) framework called DREAM (DiscoveRing and Explaining AI Mechanisms) for interpreting how AI models trained on medical images predict protected attributes. Focusing on two modalities, radiology and dermatology, we are successfully able to train high-performing classifiers for predicting race from chest x-rays (ROC-AUC score of [~]0.96) and sex from dermoscopic lesions (ROC-AUC score of [~]0.78). We highlight how incorrect use of these demographic shortcuts can have a detrimental effect on the performance of a clinically relevant downstream task like disease diagnosis under a domain shift. Further, we employ various XAI techniques to identify specific signals which can be leveraged to predict sex. Finally, we propose a technique, which we callremoval via balancing, to quantify how much a signal contributes to the classification performance. Using this technique and the signals identified, we are able to explain [~]15% of the total performance for radiology and [~]42% of the total performance for dermatology. We envision DREAM to be broadly applicable to other modalities and demographic attributes. This analysis not only underscores the importance of cautious AI application in healthcare but also opens avenues for improving the transparency and reliability of AI-driven diagnostic tools.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Spatial Transcriptomics Inferred from Pathology Whole-Slide Images Links Tumor Heterogeneity to Survival in Breast and Lung Cancer 94%
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 93%
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 93%
Similar papers in this journal
- Fostering transparent medical image AI via an image-text foundation model grounded in medical literature 97%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 92%
- Optimization and performance analytics of global aircraft-based wastewater surveillance networks 89%
Similar papers in this journal
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 95%
- Deep Learning Classification of Lipid Droplets in Quantitative Phase Images 94%
- Small hand-designed convolutional neural networks outperform transfer learning in automated cell shape detection in confluent tissues 93%
Similar papers in this journal
- Generative AI Enables Medical Image Segmentation in Ultra Low-Data Regimes 93%
- ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images 92%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.