Single-Axis Fairness Interventions Produce Asymmetric Cross-Axis Effects in Clinical Prediction
Yoon, K. Y.; Kwak, H.
Show abstract
Objective: To systematically evaluate whether single-axis demographic balancing introduces cross-axis fairness trade-offs in clinical prediction, and to characterize the directional asymmetry of these effects across model architectures and balancing strategies. Materials and Methods: We evaluated cross-axis fairness effects of demographic balancing in one-year all-cause mortality prediction using the MIMIC-IV database (N=64,427). Seven machine learning architectures were trained under six balancing strategies targeting either gender or race, with performance assessed via outcome-stratified 80-20 train-test splits repeated 30 times across both targeted and non-targeted axes using AUC, TPR, and Brier score. Results: Gender-targeting interventions largely preserved race fairness, while race-targeting consistently disrupted gender fairness across all methods and the majority of architectures. This asymmetry was invisible to same-axis evaluation alone. Race-targeting also incurred greater performance costs and calibration loss, with observed fairness gains potentially reflecting leveling down rather than genuine improvement. The same intervention could appear successful under TPR but fail under AUC evaluation. Discussion: The asymmetry likely reflects differential category complexity: binary gender balancing requires modest distributional shifts, whereas multi-category race balancing necessitates aggressive reweighting that propagates to correlated axes. Cross-axis fairness effects are directionally dependent and metric-sensitive, indicating that single-metric, single-axis evaluation is insufficient. Conclusion: Single-axis fairness optimization cannot guarantee cross-dimensional equity. Cross-axis, multi-metric fairness evaluation should be integrated into pre-deployment auditing of healthcare artificial intelligence (AI) models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 92%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 92%
Similar papers in this journal
- Personalizing renal replacement therapy initiation in the intensive care unit: a reinforcement learning-based strategy with external validation on the AKIKI randomized controlled trials 92%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 92%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 91%
Similar papers in this journal
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 93%
- Investigating Ethical Tradeoffs in Crisis Standards of Care through Simulation of Ventilator Allocation Protocols 93%
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 92%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 93%
- Predicting bloodstream infection outcome using machine learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.