Integrating Group and Individual Fairness in Clinical AI: A Post-Hoc, Model-Agnostic Framework for Fairness Auditing
Xu, J.; Hwang, Y. M.; Dormoy, I.; Jing, S. L.; Pillai, M.; Curtin, C. M.; Hernandez-Boussard, T.
Show abstract
Ensuring fairness across diverse patient populations is a fundamental challenge for clinical AI systems, yet current fairness evaluation approaches create critical blind spots. Group-level metrics capture systemic disparities but miss patient-level variations, while individual fairness frameworks ensure consistency but potentially obscure structural biases. In this paper, we propose EquiLense, a post-hoc, model-agnostic framework that bridges these perspectives through clinical similarity matching and comprehensive fairness auditing. Our method introduces the Mean Predicted Probability Difference (MPPD), which quantifies prediction inconsistencies between clinically similar patients across demographic groups, integrating both individual-level consistency and group-level equity assessment. Moreover, we provide flexible similarity matching using clinical features and comprehensive visualization tools that support practical deployment in healthcare settings. Applied to electronic health record data from over 59,000 surgical patients, our framework revealed disparities in prediction consistency even when overall model performance appeared strong. EquiLense identified differences in predicted probabilities between clinically similar patients from different racial groups, disparities that were substantially reduced when sensitive attributes were excluded from model training. Our method provides a clinically relevant and interpretable approach to fairness auditing that enables healthcare practitioners to identify, understand, and address algorithmic disparities in real-world deployment settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 94%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 95%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.