Deployment-readiness audit of calibration, clinical utility, and fairness in perioperative infection prediction
Guillen-Ramirez, H.; Lucas, K. L.; Wintsch, Y. M.; Blatter, T. U.; Triep, K.; Endrich, O.; Beldi, G.
Show abstract
Objective: Clinical risk scores intended to guide patient-level decisions can show strong average performance. However, predicted probabilities can be systematically too high or too low in specific subgroups even when overall performance is strong. We audited deployment readiness of a strong end-of-surgery postoperative infection model across clinically relevant subgroups and tested mitigation strategies in miscalibrated subgroups. Materials and Methods: We analyzed out-of-fold predictions for 10,719 surgical procedures at a Swiss tertiary hospital, with 504 postoperative bacterial infection events. Prespecified axes were recorded sex, age stratum, and an EHR-derived physiological-reserve proxy. Within subgroups and pairwise intersections, we evaluated discrimination, calibration, threshold-specific errors, and decision-curve net benefit at the prespecified operating threshold. We compared group-specific isotonic recalibration with Wasserstein-barycenter postprocessing and demonstrated portability in SUPPORT2. Results: Overall AUROC was 0.876. While sex-marginal discrimination was similar in women and men (0.878 vs 0.875), age and reserve stratification revealed deployment-readiness failures. Calibration-in-the-large ranged from -0.86 in frail patients to -2.47 in non-frail patients. At the 0.10 operating threshold, decision-curve net benefit was positive in frail patients but negative in pre-frail and non-frail patients. Isotonic recalibration corrected average physiological-reserve-stratified calibration without worsening Brier scores, whereas Wasserstein postprocessing worsened calibration in most procedure clusters. Discussion: Discrimination-only or sex-marginal evaluation would have missed subgroup failures with clinical-utility implications. Conclusion: Subgroup fairness audits for clinical deployment should jointly evaluate discrimination, calibration, and utility. We implemented the audit as the open-source isitfair framework for identifying deployment-relevant subgroup failures, comparing mitigation strategies, and generating structured reports.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 93%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 92%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 92%
Similar papers in this journal
- Ten months of temporal variation in the clinical journey of hospitalised patients with COVID-19: an observational cohort 91%
- Extent, impact, and mitigation of batch effects in tumor biomarker studies using tissue microarrays 89%
- Evaluating the effectiveness of rapid SARS-CoV-2 genome sequencing in supporting infection control teams: the COG-UK hospital-onset COVID-19 infection study 88%
Similar papers in this journal
- Outcomes of a Smartphone-based Application with Live Health-Coaching Post-Percutaneous Coronary Intervention 90%
- Older biological age is associated with adverse COVID-19 outcomes: A cohort study in UK Biobank 88%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 88%
Similar papers in this journal
- Using patient biomarker time series to determine mortality risk in hospitalised COVID-19 patients: a comparative analysis across two New York hospitals 91%
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 91%
- Determinants of hospital outcomes for COVID-19 infections in a large Pennsylvania Health System 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.