Back

Machine learning models to improve targeting of blood culture testing

Forrest-Hammond, R. W.; Gupta, R.; McVean, G.; Noursadeghi, M.; O'Grady, J.; Samuels, T. H.; Eyre, D. W.

2026-07-20 infectious diseases
10.64898/2026.07.17.26358320 medRxiv
Show abstract

Background Bloodstream infections are a major cause of mortality, yet the primary testing method, blood cultures, have low positivity (<10%) and turnaround times of 24 - 48 hours. Many are taken from patients at low risk of infection, while some bloodstream infections are diagnosed late or missed entirely. We aimed to develop and externally validate machine learning models to improve targeting of blood culture testing. Methods In this retrospective cohort study, we used routinely collected clinical and laboratory data available around culture collection from a large multi-site NHS trust (Oxford University Hospitals; Infections in Oxfordshire Research Database), between 1 January 2016 and 17 March 2025. All blood cultures taken from adults and children were included. XGBoost models were trained to predict pathogenic blood culture positivity using a temporal split (training before 1 January 2024; held-out test thereafter). External validation used emergency department data (between 1st May 2019 and 30th April 2024) from University College London Hospitals. An additional analysis examined blood culture reallocation towards the highest-risk untested admissions. Findings 294,064 cultures were included (positivity 5.6%). In the temporal hold-out test set (n=46,339), AUROC (Area Under the Receiver Operating Characteristic) was 0.853 (95% CI 0.846 - 0.860), rising to 0.876 in emergency department patients, and the model was well calibrated (slope 1.046). In external validation (n=37,326), AUROC was 0.847 (95% CI 0.839 - 0.856) with preserved calibration. In a simulated resource-neutral reallocation, replacing the 10,000 lowest-risk sent cultures with the highest-risk untested emergency admissions yielded 627 additional positive cultures (28.3% relative increase in yield). Performance was reduced when restricted to data available at the point of culture collection (AUROC 0.769, 95% CI 0.760 - 0.779). Interpretation An externally validated, well calibrated machine learning model built from broadly available, routinely collected data could improve blood culture yield without increasing testing volume, supporting resource-neutral diagnostic stewardship across NHS sites.

Matching journals

The top 14 journals account for 50% of the predicted probability mass.

1
Journal of Infection
78 papers in training set
Top 0.1%
7.8%
2
The Lancet Microbe
44 papers in training set
Top 0.1%
5.4%
3
PLOS ONE
5266 papers in training set
Top 29%
5.4%
4
Wellcome Open Research
67 papers in training set
Top 0.2%
4.3%
5
eBioMedicine
183 papers in training set
Top 0.6%
4.0%
6
Clinical Infectious Diseases
235 papers in training set
Top 0.8%
3.5%
7
BMC Infectious Diseases
133 papers in training set
Top 1%
3.2%
8
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
9
npj Digital Medicine
118 papers in training set
Top 2%
3.2%
10
The Journal of Infectious Diseases
202 papers in training set
Top 1%
2.6%
11
BMJ Open
601 papers in training set
Top 8%
2.4%
12
Communications Medicine
113 papers in training set
Top 2%
2.1%
13
Scientific Reports
3612 papers in training set
Top 48%
2.1%
14
Nature Microbiology
155 papers in training set
Top 2%
2.1%
50% of probability mass above
15
BMC Medicine
176 papers in training set
Top 2%
2.1%
16
Nature Communications
5641 papers in training set
Top 43%
2.0%
17
Eurosurveillance
83 papers in training set
Top 0.4%
1.9%
18
The Lancet Digital Health
25 papers in training set
Top 0.3%
1.7%
19
Microbial Genomics
225 papers in training set
Top 2%
1.7%
20
eLife
5828 papers in training set
Top 49%
1.7%
21
Epidemics
116 papers in training set
Top 1%
1.7%
22
Emerging Infectious Diseases
105 papers in training set
Top 0.9%
1.7%
23
Open Forum Infectious Diseases
142 papers in training set
Top 2%
1.5%
24
Frontiers in Medicine
120 papers in training set
Top 2%
1.5%
25
Journal of Clinical Pathology
15 papers in training set
Top 0.3%
1.5%
26
Microbiology Spectrum
469 papers in training set
Top 8%
1.4%
27
PLOS Biology
486 papers in training set
Top 6%
1.4%
28
PLOS Medicine
110 papers in training set
Top 2%
1.3%
29
JAMA Network Open
130 papers in training set
Top 3%
1.3%
30
American Journal of Epidemiology
67 papers in training set
Top 1%
1.0%