A universal translator for AI scores: Providing context using error
Chung, M.; Bernstein, M. H.; Yala, A.; Baird, G. L.
Show abstract
Artificial intelligence (AI) programs in radiology typically provide a numeric score for each case that correlates with the underlying pathology. However, these scores are not readily interpretable by themselves. To address this, we propose improving score interpretability by providing the False Discovery Rate (FDR) and False Omission Rate (FOR) corresponding with each score threshold. Using an open-source AI program for breast cancer, we estimated FDR and FOR across a range of AI scores using data from 130,712 digital screening mammograms, of which 907 were positive and 129,805 were negative. FDR and FOR ranged from 99.27% and 0.03%, respectively, at the low end of the score distribution to 60.98% and 0.65%, respectively, at the high end of the distribution. Providing these error rates alongside AI scores allows clinicians to consider the balance of trade-offs between false positive and false negative interpretations.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 94%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 91%
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 90%
Similar papers in this journal
- Natural language inference for clinical registry curation 92%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 91%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 90%
Similar papers in this journal
Similar papers in this journal
- Model uncertainty estimates for deep learning mammographic density prediction using ordinal and classification approaches 94%
- Breast density prediction from low and standard dose mammograms using deep learning: effect of image resolution and model training approach on prediction quality 91%
- Mammographic density assessed using deep learning in women at high risk of developing breast cancer: the effect of weight change on density 91%
Similar papers in this journal
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 92%
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 91%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.