Back

Machine learning combining FIT with up to 1,025 clinical variables: limited referral reduction but potential for faster diagnosis

Tamm, A.; Shine, B.; James, T.; Withers, J.; Salih, H.; East, J. E.; Oke, J.; Davies, J.; Morris, E. J.; Nicholson, B. D.

2026-07-17 primary care research
10.64898/2026.07.16.26358105 medRxiv
Show abstract

Background The faecal immunochemical test (FIT) is central to triaging symptomatic patients with suspected colorectal cancer (CRC) in UK primary care, yet only about one in eleven patients above the NICE 10 ug/g threshold have CRC. Existing prediction models attempting to improve on FIT have relied on conventional statistics and limited predictors. Methods GP-requested FITs with linked data (Jan 2017 - May 2025) were extracted from the Oxford University Hospitals (OUH) datawarehouse. Patients aged [≥]18 with core bloods and 180-day CRC follow-up were included. Machine learning (ML) models were trained on up to 1,025 predictors: FIT, age, sex, blood tests and their time series slopes, diagnoses/procedures/prescriptions, deprivation, BMI, and ethnicity. Models comprised penalised logistic regression, generalised additive models (EBM, NAM, SNAM, NODE-GAM), decision tree ensembles (random forests, XGBoost), and a multilayer perceptron. Referral reduction versus FIT [≥]10 ug/g was evaluated at model risk score thresholds capturing the same cancers (conservative) or same proportion of cancers (less conservative) as FIT. Potential to prioritise referred patients was assessed by examining whether positive predictive value (PPV) is very high (>30%) at any substantial sensitivity (>10%). Nested twice-repeated five-fold cross-validation provided unbiased estimates. An existing COLOFIT model was evaluated alongside. Findings 62,219 individuals (746 CRC) were analysed; 30,862 patients (315 CRC) with high/low risk symptoms and buffered FITs formed the primary subset. At [≥]10 ug/g, FIT had 91.4% sensitivity, 84.2% specificity, 5.6% PPV, and 99.9% NPV. No model reduced referrals when required to capture the same cancers as in the FIT [≥]10 ug/g cohort. Generalised additive models achieved up to 18.5% referral reduction when detecting the same proportion but some different cancers as FIT [≥]10 ug/g (EBM: 18.5%, NODE-GAM: 17.5%, SNAM: 17.4%, COLOFIT: 16.7%). At 30% sensitivity, EBM, NAM and NODE-GAM had average PPVs between 34.6%-35.0%, while FIT had a PPV of 14.6%. Interpretation Generalised additive models (GAMs) reduced referrals on average by 19% if a small proportion of the FIT-positive CRCs were substituted with originally FIT-negative CRCs by the models. No model, including COLOFIT, reduced referrals while capturing all FIT-positive cancers. Generalised additive models could detect about a third of CRCs faster, as one in three patients flagged by the models had CRC at 30% sensitivity. Funding EPSRC Centre for Doctoral Training in Health Data Science; National Institute for Health Research (NIHR) Oxford Biomedical Research Centre; Cancer Research UK. Keywords Colorectal cancer, faecal immunochemical test, machine learning, positive predictive value

Matching journals

The top 2 journals account for 50% of the predicted probability mass.