Unmeasured but Not Unbiased: The Missingness Demographic Leakage Audit (MDLA) for Calibration-Aware Fairness Evaluation in Critical Care Mortality Prediction
Patel, K.; Beedala, P.
Show abstract
ObjectiveClinical prediction models trained on electronic health records are routinely evaluated for fairness on observed feature values, but the informativeness of which measurements are absent remains unaudited. We developed the Missingness Demographic Leakage Audit (MDLA), a reproducible four-step informatics framework that tests whether patterns of clinical measurement absence function as latent demographic proxies -- constituting a bias pathway invisible to standard fairness audits. Materials and MethodsWe applied MDLA across development (MIMIC-IV v2.2; n=50,827; mortality 10.2%) and external validation (eICU-CRD v2.0; n=137,773; mortality 9.5%) cohorts following TRIPOD+AI standards. XGBoost, random forest, and logistic regression were trained on 43 clinical features and 44 binary missingness indicators. MDLA quantified demographic predictability from missingness alone, tested feature-level associations with Bonferroni correction, and verified model reliance via ablation. A calibration-aware fairness audit evaluated five criteria across four demographic axes; six post-hoc recalibration strategies were compared on a fairness-utility Pareto frontier. ResultsMissingness indicators alone predicted racial group membership above chance (AUROC=0.543; 95% CI, 0.540-0.546), with 18 of 43 features showing Bonferroni-significant race-missingness associations (all Cramers V<0.10). Ablation confirmed model reliance: adding missingness indicators increased racial AUROC disparity by 10.7% (0.063 to 0.069) without improving global performance. XGBoost achieved AUROC=0.910 internally (AUROC=0.799 on external validation). Global Platt recalibration reduced overall calibration error by 94% and maximum racial calibration error by 51%, with zero AUROC loss and successful parameter transfer to external validation without retraining. ConclusionMDLA provides a structured, reproducible protocol for detecting missingness-encoded demographic signals prior to model deployment. Applied across 188,600 ICU patient-stays from two institutionally diverse databases, it identified a statistically confirmed but subtle bias pathway undetectable by standard fairness audits. Missingness-aware auditing and calibration-aware evaluation should be integrated into clinical AI validation pipelines.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 95%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
Similar papers in this journal
- GLUCOSE: A Distributional Reinforcement Learning Model for Optimal Glucose Control After Cardiac Surgery 94%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 93%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 93%
Similar papers in this journal
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 94%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 94%
- Using patient biomarker time series to determine mortality risk in hospitalised COVID-19 patients: a comparative analysis across two New York hospitals 93%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Personalizing renal replacement therapy initiation in the intensive care unit: a reinforcement learning-based strategy with external validation on the AKIKI randomized controlled trials 94%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 93%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 95%
- Predicting bloodstream infection outcome using machine learning 95%
- Developing And Validating COVID-19 Adverse Outcome Risk Prediction Models From A Bi-National European Cohort Of 5594 Patients 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.