Large-scale evaluation of an AI system as an independent reader for double reading in breast cancer screening
Sharma, N.; Ng, A. Y.; James, J. J.; Khara, G.; Ambrozay, E.; Austin, C. C.; Forrai, G.; Glocker, B.; Heindl, A.; Karpati, E.; Rijken, T. M.; Venkataraman, V.; Yearsley, J. E.; Kecskemethy, P. D.
Show abstract
ImportanceScreening mammography with two human readers increases cancer detection and lowers recall rates, but workforce shortages make double reading unsustainable in many countries. Artificial intelligence (AI) as an independent reader in double reading may support screening performance while improving cost-effectiveness. The clinical validation of AI requires large-scale, multi-vendor studies on unenriched cohorts. ObjectiveTo evaluate the performance of the Mia(R) AI system on data that the AI system would process in real-world deployments. DesignA retrospective study simulating the impact of AI on an unenriched screening sample. SettingSeven European breast screening sites representing four centers: three from the UK and one in Hungary (HU), between 2009 and 2019. ParticipantsThe sample included 275,900 cases (177,882 participants) from seven screening sites, involving two countries and four hardware vendors from 2009 to 2019. InterventionSimulation of double reading using AI as an independent reader in breast cancer screening on historical data. Main Outcomes and MeasuresPerformance was determined for standalone AI compared to the historical single reader and for simulated double reading with AI compared to historical double reading, assessing non-inferiority and superiority on relevant screening metrics using a non-inferiority margin of 10% relative difference and a one-sided alpha of 2.5% for both tests. ResultsStandalone AI detected 29.8% of missed interval cancers. When compared with historical double reading, double reading with AI showed non-inferiority for sensitivity and superiority for recall rate, specificity and positive predictive value. AI as an independent reader reduced the workload for the second human reader but increased the arbitration rate from 3.3% to 12.3%. Applying the AI system could have reduced the human reading time required by up to 44.8% and reduced the recall rate by a relative 7.7% (from 5.2% to 4.8%). Conclusions and RelevanceUsing the AI system as an independent reader maintains or improves the double reading standard of care, while substantially reducing the workload. Thus, it has the potential to provide operational and economic benefits. Trial RegistrationRegistered on ISRCTN, study ID: ISRCTN18056078
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 94%
- BREAst screening Tailored for HEr (BREATHE) - A Study Protocol On Personalised Risk-based Breast Cancer Screening Programme 92%
- Understanding motivations of older women to continue or discontinue breast cancer screening 91%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 89%
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 89%
Similar papers in this journal
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 94%
- The NILS study protocol - a retrospective validation study of a preoperative decision-making tool for non-invasive lymph node staging in women with primary breast cancer [ISRCTN14341750] 93%
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 92%
Similar papers in this journal
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 94%
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 90%
- An economic evaluation of two self-sampling strategies for HPV primary cervical cancer screening compared with clinician-collected sampling 90%
Similar papers in this journal
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 94%
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 93%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.