Assessing eligibility for lung cancer screening: Parsimonious multi-country ensemble machine learning models for lung cancer prediction
Callender, T.; Imrie, F.; Cebere, B.; Pashayan, N.; Navani, N.; van der Schaar, M.; Janes, S. M.
Show abstract
BackgroundEnsemble machine learning could support the development of highly parsimonious prediction models that maintain the performance of more complex models whilst maximising simplicity and generalisability, supporting the widespread adoption of personalised screening. In this work, we aimed to develop and validate ensemble machine learning models to determine eligibility for risk-based lung cancer screening. MethodsFor model development, we used data from 216,714 ever-smokers in the UK Biobank prospective cohort and 26,616 high-risk ever-smokers in the control arm of the US National Lung Screening randomised controlled trial. We externally validated our models amongst the 49,593 participants in the chest radiography arm and amongst all 80,659 ever-smoking participants in the US Prostate, Lung, Colorectal and Ovarian Screening Trial (PLCO). Models were developed to predict the risk of two outcomes within five years from baseline: diagnosis of lung cancer, and death from lung cancer. We assessed model discrimination (area under the receiver operating curve, AUC), calibration (calibration curves and expected/observed ratio), overall performance (Brier scores), and net benefit with decision curve analysis. ResultsModels predicting lung cancer death (UCL-D) and incidence (UCL-I) using three variables - age, smoking duration, and pack-years - achieved or exceeded parity in discrimination, overall performance, and net benefit with comparators currently in use, despite requiring only one-quarter of the predictors. In external validation in the PLCO trial, UCL-D had an AUC of 0.803 (95% CI: 0.783-0.824) and was well calibrated with an expected/observed (E/O) ratio of 1.05 (95% CI: 0.95-1.19). UCL-I had an AUC of 0.787 (95% CI: 0.771-0.802), an E/O ratio of 1.0 (0.92-1.07). The sensitivity of UCL-D was 85.5% and UCL-I was 83.9%, at 5-year risk thresholds of 0.68% and 1.17%, respectively 7.9% and 6.2% higher than the USPSTF-2021 criteria at the same specificity. ConclusionsWe present parsimonious ensemble machine learning models to predict the risk of lung cancer in ever-smokers, demonstrating a novel approach that could simplify the implementation of risk-based lung cancer screening in multiple settings.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 91%
- Predictive performance and clinical application of COV50, a urinary proteomic biomarker in early COVID-19 infection: a cohort study 90%
- Predicting hospital-onset COVID-19 infections using dynamic networks of patient contacts: an observational study 90%
Similar papers in this journal
- Radiomics analysis to predict pulmonary nodule malignancy using machine learning approaches 93%
- The protective effect of club cell secretory protein (CC-16) on COPD risk and progression: a Mendelian randomisation study 92%
- Post-viral parenchymal lung disease following COVID-19 and viral pneumonitis hospitalisation: A systematic review and meta-analysis 91%
Similar papers in this journal
- Genetic associations and architecture of asthma-chronic obstructive pulmonary disease overlap 91%
- Rethinking Blood Eosinophils for Assessing ICS Response in COPD: A Post-Hoc Analysis from FLAME 89%
- Immediate, remote smoking cessation intervention in participants undergoing a targeted lung health check: QuLIT2 a randomised controlled trial 88%
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 90%
- DeePaN: A deep patient graph convolutional network integrating clinico-genomic evidence to stratify lung cancers benefiting from immunotherapy 90%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 90%
Similar papers in this journal
- Integration of clinical characteristics, lab tests and a deep learning CT scan analysis to predict severity of hospitalized COVID-19 patients 93%
- Prediction and stratification of longitudinal risk for chronic obstructive pulmonary disease across smoking behaviors 92%
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.