Back

Prediction of Hospital Outpatient Attendance in UK Hospitals: A Retrospective Study Applying Machine Learning to Routinely Collected Data for Patients of All Ages

Holdship, J.; Dhanoa, H.; Hopper, A.; Steves, C. J.; Butler, M.; Wolfe, I.; Tucker, K.; Cooper, C.; Yates, J.

2022-05-31 health informatics
10.1101/2022.01.24.22269733 medRxiv
Show abstract

ObjectivesPatient non-attendance at outpatient appointments is a major concern for healthcare providers. Non-attendances increase waiting lists, reduce access to care and may be detrimental not for the patient who did not attend. We aim to produce a model which can accurately predict which appointments will be attended. SettingA teaching hospital in London, UK combining secondary and tertiary care. ParticipantsA set of 9.6 million outpatient appointments between April 2015 and September 2019 including all ages and specialities. Primary and secondary outcome measuresArea under the receiver operating characteristic curve (AU-ROC) for prediction of outpatient appointment non-attendances. ResultsThe model uses 27 predictors to achieve an AUROC score of 0.768 (95% CI: 0.767-0.769) and accuracy of 89.2% (95% CI: 89.16%-89.24%) on test data. We find that the waiting period between booking and the appointment, the patients past attendance behaviour, and the levels of deprivation in their local area are important factors in predicting future attendance. ConclusionOur model successfully predicts patient attendance at outpatient appointments. Its performance on both patients who did not appear in the training data and appointments from a different time period which covers the Covid-19 pandemic indicate it generalized well across both face to face and virtual appointments and could be used to target resources and intervention towards those patients who are likely to miss an appointment. Moreover, it highlights the impact of deprivation on patient access to healthcare Strengths and Limitation of this StudyO_LIWe make use of a large dataset which enables us to use complex machine learning algorithms. C_LIO_LIWe validate the model on two large, distinct datasets giving high confidence in our model performance. C_LIO_LIAn unknown amount of patient data is missing due to a nearby hospital which shares patients with the study setting. C_LI

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.