A rapid review exploring the effectiveness of artificial intelligence for cancer diagnosis
Wale, A.; Shaw, H.; Ayres, T.; Okolie, C.; Edwards, R. T.; Davies, J.; Lewis, R.; Cooper, A.; Edwards, A. G.
Show abstract
There is growing demand for diagnostic services in the UK. This rapid review aimed to assess the effectiveness of artificial intelligence (AI) in diagnostic radiology with a focus on cancer diagnosis. A range of AI models including machine learning, deep learning and ensemble models, were assessed in this review. The review included an initial broad mapping exercise and a more in-depth synthesis of a specific sub-set of the evidence. The review included evidence available from 2018 until June 2023. A total of 92 comparative primary studies were included in the evidence map. The evidence map identified 52 studies in which the AI models were in the early stages of development and validation, and highlighted breast, lung and prostate cancers as the type of cancers most frequently reported on. 28 studies evaluating an established model and focusing on the diagnosis of breast, lung, and prostate cancer were included in the in-depth synthesis. All studies included in the in-depth synthesis were classified as diagnostic accuracy studies. Only one study evaluated an AI model that was commercially available in the UK. Most studies reported results in favour of the AI models, however, these improvements were not always statistically significant. The studies also varied considerably in terms of AI models studied, type of cancer, images used, and comparison made; and were limited in terms of their methodology. When used as a standalone diagnostic tool, there is evidence to suggest that AI can improve diagnostic accuracy or is comparable to experienced radiologists, however this may be dependent on the AI model being used. There is evidence to suggest that AI may be beneficial when used as a support tool for clinicians/radiologists with less experience. The impact of AI on the timeline involved in diagnosis appeared inconsistent. AI may speed up the diagnostic timeline when the level of cancer suspicion is low but may increase diagnostic timelines when the level of cancer suspicion is high. The evidence suggests that clinicians are accepting of AI-based assistance for cancer diagnosis. Policy and practice implicationsThe overall evidence for effectiveness appeared in favour of AI and several factors were identified that impact the effectiveness of the AI models. AI may improve diagnostic accuracy in clinicians/radiologists with less experience of interpreting radiological images. However, further well-designed high-quality research is needed from the UK and similar countries to better understand the effectiveness of AI in cancer diagnosis. Economic considerationsThere is little evidence on the cost-effectiveness of using AI for cancer diagnosis. In theory, it might be possible for AI to assist with earlier diagnosis of cancer with both health and economic benefits. Funding statementThe Public Health Wales Observatory was funded for this work by the Health and Care Research Wales Evidence Centre, itself funded by Health and Care Research Wales on behalf of Welsh Government. EXECUTIVE SUMMARYO_ST_ABSWhat is a Rapid Review?C_ST_ABSOur rapid reviews (RR) use a variation of the systematic review approach, abbreviating or omitting some components to generate the evidence to inform stakeholders promptly whilst maintaining attention to bias. Who is this Rapid Review for?The review question was suggested by the Health Sciences Directorate (Policy). Background / Aim of Rapid ReviewThere is growing demand for diagnostic services in the UK. The use of artificial intelligence in diagnosis is part of the Welsh Governments programme for transforming and modernising planned care and reducing waiting lists in Wales. This rapid review aimed to assess the effectiveness of artificial intelligence (AI) in diagnostic radiology with a focus on cancer diagnosis. A range of AI models including machine learning, deep learning and ensemble models, were assessed in this review. The term AI models was therefore used to encompass these different types of AI models described in the literature. The review included an initial broad mapping exercise and a more in-depth synthesis of a specific sub-set of the evidence. The focus of the in-depth synthesis was informed by the reviews stakeholders based on the findings of the mapping exercise. ResultsO_ST_ABSRecency of the evidence baseC_ST_ABSO_LIThe review included evidence available from 2018 until June 2023. C_LI Extent of the evidence baseO_LIA total of 92 comparative primary studies were included in the evidence map. C_LIO_LIThe evidence map identified 52 studies in which the AI models were in the early stages of development and validation, and highlighted breast, lung and prostate cancers as the type of cancers most frequently reported on. C_LIO_LI28 studies evaluating an established model and focusing on the diagnosis of breast (n=14), lung (n=7) and prostate (n=7) cancer were included in the in-depth synthesis. C_LIO_LIStudies included in the in-depth synthesis were conducted in the USA (n=8), Japan (n=5), UK (n=2), Italy (n=2), Turkey (n=2), Germany (n=2), Netherlands (n=2), Portugal (n=1), Greece (n=1) and Norway (n=1). Two studies were conducted across multiple countries. C_LIO_LIAll studies included in the in-depth synthesis were classified as diagnostic accuracy studies. C_LIO_LIOnly one study evaluated an AI model that was commercially available in the UK. C_LIO_LIA total of 14 studies compared AI models to human readers or to other diagnostic methods used in practice, 13 studies compared the impact of AI on human interpretation of radiologic images when diagnosing cancer, four studies compared multiple AI models, and one study compared an inexperienced AI-assisted reader with an experienced reader without AI. C_LIO_LIFive studies reported on the impact of AI on diagnostic timelines (time to diagnosis, assessment time, evaluation times, and reading time). C_LIO_LIFour studies also reported on the impact of AI on inter/intra-reader variability, reliability, and agreement. C_LIO_LIOne study reported on clinicians acceptance and receptiveness of the use of AI for cancer diagnosis. C_LI Key findings and certainty of the evidenceO_LIMost studies reported results in favour of the AI models, however, these improvements were not always statistically significant. The studies also varied considerably in terms of AI models studied, type of cancer, images used, and comparison made; and were limited in terms of their methodology (unclear level of certainty). C_LIO_LIWhen used as a standalone diagnostic tool, there is evidence to suggest that AI can improve diagnostic accuracy or is comparable to experienced radiologists, however this may be dependent on the AI model being used (unclear level of certainty). C_LIO_LIThere is evidence to suggest that AI may be beneficial when used as a support tool for clinicians/radiologists with less experience (unclear level of certainty). C_LIO_LIThe impact of AI on the timeline involved in diagnosis appeared inconsistent. AI may speed up the diagnostic timeline when the level of cancer suspicion is low but may increase diagnostic timelines when the level of cancer suspicion is high (low level of certainty). C_LIO_LIThe evidence suggests that clinicians are accepting of AI-based assistance for cancer diagnosis (low level of certainty). C_LI Research Implications and Evidence GapsO_LINo study reported on any patient outcomes, including patient harms. C_LIO_LINo study reported on any economic outcomes. C_LIO_LINo study reported on equity outcomes, including equity of access. C_LIO_LIFurther research in a real-world setting is needed to better understand the cost implications and impact on patient safety of AI for cancer diagnosis. C_LI Policy and Practice ImplicationsO_LIThe overall evidence for effectiveness appeared in favour of AI and several factors were identified that impact the effectiveness of the AI models. C_LIO_LIAI may improve diagnostic accuracy in clinicians/radiologists with less experience of interpreting radiological images. C_LIO_LIAI models are continually being developed and updated and findings are likely to vary between different AI models. C_LIO_LIFurther well-designed high-quality research is needed from the UK and similar countries to better understand the effectiveness of AI in cancer diagnosis. C_LI Economic considerationsO_LIIn theory it might be possible for AI to assist with earlier diagnosis of cancer with both health and economic benefits. C_LIO_LIThere is little evidence on the cost-effectiveness of using AI for cancer diagnosis. One modelling paper from the United States (US) suggests using AI in lung cancer screening using low-dose computerised tomography (CT) scans can be cost-effective, up to a cost of $1,240 per patient screened. C_LIO_LIThe UK (and its constituent countries) perform consistently poorly against European and international comparators in terms of cancer survival rates. Cancer screening was suspended and routine diagnostic work deferred in the UK during the COVID-19 pandemic. C_LIO_LIThe cost of cancer to the UK economy in 2019 was estimated to be least {pound}1.4 billion a year in lost wages and benefits alone. When widening the perspective to include mortality, this figure rises to {pound}7.6 billion a year. Pro-rating both figures to the Welsh economy and adjusting for inflation gives figures of {pound}79 million and {pound}429 million per annum respectively C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 92%
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 92%
- Surgical Resection, Radiotherapy, And Percutaneous Thermal Ablation for Treatment of Stage 1 Non-Small Cell Lung Cancer: A Systematic Review and Network Meta-Analysis 91%
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 94%
- National diagnostic reference levels for digital diagnostic and screening mammography in Uganda. 92%
- Early user experience and lessons learned using ultra-portable digital X-ray with computer-aided detection (DXR-CAD) products: A qualitative study from the perspective of healthcare providers 92%
Similar papers in this journal
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 90%
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 90%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.