Evaluating and Reducing Subgroup Disparity in AI Models: An Analysis of Pediatric COVID-19 Test Outcomes
Libin, A.; Treitler, J. T.; Vasaitis, T.; Shao, Y.
Show abstract
Artificial Intelligence (AI) fairness in healthcare settings has attracted significant attention due to the concerns to propagate existing health disparities. Despite ongoing research, the frequency and extent of subgroup fairness have not been sufficiently studied. In this study, we extracted a nationally representative pediatric dataset (ages 0-17, n=9,935) from the US National Health Interview Survey (NHIS) concerning COVID-19 test outcomes. For subgroup disparity assessment, we trained 50 models using five machine learning algorithms. We assessed the models area under the curve (AUC) on 12 small (<15% of the total n) subgroups defined using social economic factors versus the on the overall population. Our results show that subgroup disparities were prevalent (50.7%) in the models. Subgroup AUCs were generally lower, with a mean difference of 0.01, ranging from -0.29 to +0.41. Notably, the disparities were not always statistically significant, with four out of 12 subgroups having statistically significant disparities across models. Additionally, we explored the efficacy of synthetic data in mitigating identified disparities. The introduction of synthetic data enhanced subgroup disparity in 57.7% of the models. The mean AUC disparities for models with synthetic data decreased on average by 0.03 via resampling and 0.04 via generative adverbial network methods.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
Similar papers in this journal
- A scoping review of fair machine learning techniques when using real-world data 95%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 92%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 92%
Similar papers in this journal
- Machine Learning Approaches for Electronic Health Records Phenotyping: A Methodical Review 93%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 92%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 92%
Similar papers in this journal
Similar papers in this journal
- Algorithmic Individual Fairness and Healthcare: A Scoping Review 93%
- Modeling physician variability to prioritize relevant medical record information 93%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.