An AI-Guided Data Centric Strategy to Detect and Mitigate Biases in Healthcare Datasets
Gulamali, F.; Sawant, A.; Liharska, L.; Horowitz, C. R.; Chan, L.; Kovatch, P. H.; Hofer, I.; Singh, K.; Richardson, L. D.; Mensah, E.; Charney, A.; Reich, D.; Hu, J.; Nadkarni, G.
Show abstract
BackgroundBroad adoption of artificial intelligence (AI) algorithms in healthcare has led to perpetuation of bias found in datasets used for algorithm training. Methods to mitigate bias involve approaches after training leading to tradeoffs between sensitivity and specificity. There have been limited efforts to address bias at the level of the data for algorithm generation. MethodsWe generate a data-centric, but algorithm-agnostic approach to evaluate dataset bias by investigating how the relationships between different groups are learned at different sample sizes. We name this method AEquity and define a metric AEq. We then apply a systematic analysis of AEq values across subpopulations to identify and mitigate manifestations of racial bias. FindingsWe demonstrate that AEquity helps mitigate different biases in three different chest radiograph datasets, a healthcare costs dataset, and when using tabularized electronic health record data for mortality prediction. In the healthcare costs dataset, we show that AEquity is a more sensitive metric of label bias than model performance. AEquity can be utilized for label selection when standard fairness metrics fail. In the chest radiographs dataset, we show that AEquity can help optimize dataset selection to mitigate bias, as measured by nine different fairness metrics across nine of the most frequent diagnoses and four different protected categories (race, sex, insurance status, age) and the intersections of race and sex. We benchmark against approaches currently used after algorithm training including recalibration and balanced empirical risk minimization. Finally, we utilize AEquity to characterize and mitigate a previously unreported bias in mortality prediction with the widely used National Health and Nutrition Examination Survey (NHANES) dataset, showing that AEquity outperforms currently used approaches, and is effective at both small and large sample sizes. InterpretationAEquity can identify and mitigate bias in known biased datasets through different strategies and an unreported bias in a widely used dataset. SummaryAEquity, a machine learning approach can identify and mitigate bias the level of datasets used to train algorithms. We demonstrate it can mitigate known cases of bias better than existing methods, and detect and mitigate bias that was previously unreported. EVIDENCE IN CONTEXTO_ST_ABSEvidence before this studyC_ST_ABSMethods to mitigate algorithmic bias typically involve adjustments made after training, leading to a tradeoff between sensitivity and specificity. There have been limited efforts to mitigate bias at the level of the data. Added value of this studyThis study introduces a machine learning based method, AEquity, which analyzes the learnability of data from subpopulations at different sample sizes, which can then be used to intervene on the larger dataset to mitigate bias. The study demonstrates the detection and mitigation of bias in two scenarios where bias had been previously reported. It also demonstrates the detection and mitigation of bias the widely used National Health and Nutrition Examination Survey (NHANES) dataset, which was previously unknown. Implications of all available evidenceAEquity is a complementary approach that can be used early in the algorithm lifecycle to characterize and mitigate bias and thus prevent perpetuation of algorithmic disparities.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 95%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
- Machine Learning for Real-Time Aggregated Prediction of Hospital Admission for Emergency Patients 94%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
- Community-acquired pneumonia identification from electronic health records in the absence of a gold standard: a Bayesian latent class analysis 93%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 96%
- Synthetic data for privacy-preserving clinical risk prediction 95%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 95%
Similar papers in this journal
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 93%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 92%
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.