External Validation of Predictive Models for Diagnosis, Management and Severity of Pediatric Appendicitis
Marcinkevics, R.; Sokol, K.; Paulraj, A.; Hilbert, M. A.; Rimili, V.; Wellmann, S.; Knorr, C.; Reingruber, B.; Vogt, J. E.; Reis Wolfertstetter, P.
Show abstract
BackgroundAppendicitis is a common condition among children and adolescents. Machine learning models can offer much-needed tools for improved diagnosis, severity assessment and management guidance for pediatric appendicitis. However, to be adopted in practice, such systems must be reliable, safe and robust across various medical contexts, e.g., hospitals with distinct clinical practices and patient populations. MethodsWe performed external validation of models predicting the diagnosis, management and severity of pediatric appendicitis. Trained on a cohort of 430 patients admitted to the Childrens Hospital St. Hedwig (Regensburg, Germany), the models were validated on an independent cohort of 301 patients from the Florence-Nightingale-Hospital (Dusseldorf, Germany). The data included demographic, clinical, scoring, laboratory and ultrasound parameters. In addition, we explored the benefits of model retraining and inspected variable importance. ResultsThe distributions of most parameters differed between the datasets. Consequently, we saw a decrease in predictive performance for diagnosis, management and severity across most metrics. After retraining with a portion of external data, we observed gains in performance, which, nonetheless, remained lower than in the original study. Notably, the most important variables were consistent across the datasets. ConclusionsWhile the performance of transferred models was satisfactory, it remained lower than on the original data. This study demonstrates challenges in transferring models between hospitals, especially when clinical practice and demographics differ or in the presence of externalities such as pandemics. We also highlight the limitations of retraining as a potential remedy since it could not restore predictive performance to the initial level.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An effect of the COVID-19 pandemic: significantly more complicated appendicitis due to delayed presentation of patients! 94%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 93%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 92%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Artificial Intelligence Model for Analyzing Colonic Endoscopy Images to Detect Changes Associated with Irritable Bowel Syndrome 92%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
Similar papers in this journal
Similar papers in this journal
- The impact of site-specific clearance on Methicillin-resistant Staphylococcus aureus decolonization 89%
- Mathematical modelling of the immune response during endometriosis lesion onset 89%
- A semi-parametric, state-space compartmental model with time-dependent parameters for forecasting COVID-19 cases, hospitalizations, and deaths 88%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 95%
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.