Refocusing Algorithmic Fairness on Feature-Level Bias: A Diagnostic Approach Using Dutch EHR Data
Parsons, C. S.; Girwar, S.-A. M.; Azimi, S.; Spruit, M. R.; Haas, M. R.
Show abstract
While algorithmic fairness research in healthcare has predominantly focused on disparities in model performance, less attention has been given to the underlying data structures that may drive such disparities. High-level fairness metrics often obscure the deeper feature-level dynamics necessary for a critical and context-aware assessment of fairness. To address this gap, we propose and apply a diagnostic framework termed Feature-Level Bias Identification And Sensemaking (FL-BIAS). As a case study, we conducted a secondary analysis of a retrospective cross-sectional cohort study using electronic health record (EHR) data from Dutch general practitioners, linked with sociodemographic data from Statistics Netherlands. The dataset included 112,872 patients, of whom 16.2% had a non-Western migration background. Hospitalization in the following year was modeled using Johns Hopkins Aggregated Diagnosis Groups (ADGs). We trained logistic regression and XGBoost models on different subgroup datasets to evaluate performance disparities using fairness metrics. To analyze feature-level contributions to predictions and errors, we applied Shapley Value methods, including Kernel SHAP and Cohort Shapley. Exploratory analysis revealed significant differences in SES and ADG distributions between Dutch and non-Western groups, though Multiple Correspondence Analysis showed minimal structural variation. Mediation analysis indicated that the effect of migration background on hospitalization was largely mediated by SES, with potential unobserved confounding. While standard fairness metrics indicated modest bias in favor of non-Western patients, deeper feature-level analyses revealed subgroup-specific patterns of variable importance that suggest potentially less favorable underlying conditions. For instance, malignancy (ADG 32) had a stronger predictive impact among non-Western patients but contributed less to false negatives compared to Dutch patients, possibly reflecting structural disparities in cancer diagnosis and care. These findings highlight the need for contextual, multi-level evaluations of algorithmic bias. Fairness in healthcare AI must be approached as a socio-technical challenge, requiring multidisciplinary collaboration to uncover root causes and guide effective mitigation strategies. Author SummaryUnfair algorithms often originate in the data itself. To help detect such hidden biases, we combined known data science methods into a diagnostic approach named Feature-Level Bias Identification And Sensemaking (FL-BIAS). Using Dutch general practitioner records linked with national demographic information, we explored how patient migration background corrected for socioeconomic status affect predictions of hospital admissions. We discovered that certain medical conditions, such as cancer and chronic illness, contributed differently to predictions for Dutch compared to non-Western patients, revealing subtle but important patterns in how data reflects social inequalities. Our goal was to move beyond simple fairness scores and better understand how inequalities can be hidden in health data, supporting the design of fairer and more transparent healthcare AI tools that truly serve all patients.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 95%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 94%
- A scoping review of fair machine learning techniques when using real-world data 94%
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 93%
Similar papers in this journal
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 94%
- A Multi-Granular Stacked Regression for Forecasting Long-Term Demand in Emergency Departments 93%
- Development of a data-driven COVID-19 prognostication tool to inform triage and step-down care for hospitalised patients in Hong Kong: A population based cohort study 93%
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 94%
- Temporally-Informed Random Forests for Suicide Risk Prediction 93%
Similar papers in this journal
- Latent class regression improves the predictive acuity and clinical utility of survival prognostication amongst chronic heart failure patients. 93%
- Patterns of rates of mortality in the Clinical Practice Research Datalink 93%
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.