AI integration improves breast cancer screening in a real-world, retrospective cohort study
Frazer, H. M.; Pena-Solorzano, C. A.; Kwok, C. F.; Elliott, M.; Chen, Y.; Wang, C.; the BRAIx team, ; Lippey, J.; Hopper, J.; Brotchie, P.; Carneiro, G.; McCarthy, D. J.
Show abstract
Artificial intelligence (AI) holds promise for improving breast cancer screening, but many challenges remain in implementing AI tools in clinical screening services. AI readers compare favourably against individual human radiologists in detecting breast cancer in population screening programs. However, single AI or human readers cannot perform at the level of multi-reader systems such as those used in Australia, Sweden, the UK, and other countries. The implementation of AI readers in mammographic screening programs therefore demands integration of AI readers in multi-reader systems featuring collaboration between humans and AI. Successful integration of AI readers demands a better understanding of possible models of human-AI collaboration and exploration of the range of possible outcomes engendered by the effects on human readers of interacting with AI readers. Here, we used a large, high-quality retrospective mammography dataset from Victoria, Australia to conduct detailed simulations of five plausible AI-integrated screening pathways. We compared the performance of these AI-integrated pathways against the baseline standard-of-care "two reader plus third arbitration" system used in Australia. We examined the influence of positive, neutral, and negative human-AI interaction effects of varying strength to explore possibilities for upside, automation bias, and downside risk of human-AI collaboration. Replacing the second reader or allowing the AI reader to make high confidence decisions can improve upon the standard of care screening outcomes by 1.9-2.5% in sensitivity and up to 0.6% in specificity (with 4.6-10.9% reduction in the number of assessments and 48-80.7% reduction in the number of reads). Automation bias degrades performance in multi-reader settings but improves it for single-readers. Using an AI reader to triage between single and multi-reader pathways can improve performance given positive human-AI interaction. This study provides insight into feasible approaches for implementing human-AI collaboration in population mammographic screening, incorporating human-AI interaction effects. Our study provides evidence to support the urgent assessment of AI-integrated screening pathways with prospective studies to validate real-world performance and open routes to clinical adoption.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Diagnostic Accuracy of Artificial Intelligence in Classifying HER2 Status in Breast Cancer Immunohistochemistry Slides and Implications for HER2-Low Cases: A Systematic Review and Meta-Analysis 94%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 92%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 91%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 94%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 90%
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 89%
Similar papers in this journal
- Artificial Intelligence System Reduces False-Positive Findings in the Interpretation of Breast Ultrasound Exams 95%
- Deep Learning Allows Assessment of Risk of Metastatic Relapse from Invasive Breast Cancer Histological Slides 92%
- Integrated radiogenomics models predict response to neoadjuvant chemotherapy in high grade serous ovarian cancer 91%
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- BREAst screening Tailored for HEr (BREATHE) - A Study Protocol On Personalised Risk-based Breast Cancer Screening Programme 91%
- Understanding motivations of older women to continue or discontinue breast cancer screening 90%
Similar papers in this journal
- Assessing generalizability of an AI-based visual test for cervical cancer screening 93%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 93%
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.