Triage with AI: A Rule-out Framework Quantifying the Risks and Benefits of Screening Mammogram Automation
Bernstein, M. H.; Chung, M.; Yala, A.; Baird, G. L.
Show abstract
BackgroundAI has been proposed as a triage or "rule out" device to reduce radiologist workload, but it is presently unclear how an AI triage threshold should be determined. We present a framework for determining an optimal threshold. Materials and Methods114,229 bilateral 2D digital screening mammograms were retrospectively analyzed from 2006-2023. All mammograms were given an AI score using Mirai, an open-source deep-learning model. Several metrics were examined using two thresholds for determining ruled out versus retained cases: 1) Caseload Reduce Rate (CRR; percent of caseload reduced due to rule-out), 2) Gross AI False Omission Rate (G-FOR; probability of a patient having breast cancer if ruled out), 3) AI Net False Omission Rate (N-FOR; probability of a patient having breast cancer if ruled out and the radiologist would have caught in standard care [i.e. no triage].), 4) AI Adjusted Net False Omission Rate (30%) (AN-FOR[30%]; N-FOR adjusted for the hypothetical scenario where radiologists detect an extra 30% of breast cancers among AI retained cases). The two thresholds were severity scores of 0.2 (Yudens J) and 0.05 (AN-FOR[30%]=0). The former is mathematically optimal; the latter reflects a threshold where AI triage does not introduce any total increase in False Negatives. ResultsAt the 0.20 threshold, G-FOR, N-FOR, and AN-FOR(30%) were 0.26%, 0.017%, and 0.14%, respectively (223, 141, and 121, respectively, missed cancer cases) and CRR=75%. At the 0.05 threshold, the G-FOR, N-FOR, and AN-FOR (30%) are 0.12%, 0.07%, and 0.00% (49, 30, and 0, respectively, missed cancer cases) and CRR=36%. ConclusionWe demonstrate how radiology practices can consider the trade-offs of using different AI scores triage thresholds. At the AN-FOR rate of 30%, the Yudens J threshold results in 121 additional missed cancers for a 75% caseload reduction. We estimate no additional missed cancers at a 36% caseload reduction.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 94%
- Understanding motivations of older women to continue or discontinue breast cancer screening 91%
- BREAst screening Tailored for HEr (BREATHE) - A Study Protocol On Personalised Risk-based Breast Cancer Screening Programme 91%
Similar papers in this journal
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 92%
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 92%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 91%
Similar papers in this journal
- An updated PREDICT breast cancer prognostic model including the benefits and harms of radiotherapy 91%
- Investigating the relationship between breast cancer risk factors and an AI-generated mammographic texture feature in the Nurses' Health Study II 91%
- Unmasking the tissue microecology of ductal carcinoma in situ with deep learning 89%
Similar papers in this journal
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 91%
- Incorporating Polygenic Risk Scores and Nongenetic Risk Factors for Breast Cancer Risk Prediction among Asian Women, Results from Asia Breast Cancer Consortium 90%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 88%
Similar papers in this journal
- Model uncertainty estimates for deep learning mammographic density prediction using ordinal and classification approaches 93%
- Mammographic density assessed using deep learning in women at high risk of developing breast cancer: the effect of weight change on density 92%
- Breast density prediction from low and standard dose mammograms using deep learning: effect of image resolution and model training approach on prediction quality 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.