Correcting Algorithmic Bias in Machine Learning Prediction of Healthcare utilization in India
Lee, J. T.; Li, V. C.-S.; Hsu, S. H.; Lu, T.-P.; Wang, C.; Perianayagam, A.; Anindya, K.; Atun, R.
Show abstract
ObjectiveThis study investigates how historical disparities in healthcare access influence machine learning (ML) predictions of healthcare utilization among older adults in India through algorithmic bias. We examine the extent to which standard ML models underestimate utilization in disadvantaged populations and quantify the resulting distortion in national-level cost projections. MethodsUsing data from 55,698 respondents in the Longitudinal Ageing Study in India (LASI), we trained two sets of ML models to predict outpatient and inpatient utilization: one on the full population (Model 1), and another on a subsample with met healthcare needs (Model 2). Gradient Boosting was selected as the best-performing algorithm. To interpret model predictions and identify key drivers of healthcare utilization, we applied SHapley Additive exPlanations (SHAP). We compared model outputs across socioeconomic subgroups and extrapolated predicted utilization to national population estimates using WHO-CHOICE unit costs. FindingsModel 1 consistently underestimated healthcare utilization relative to Model 2, particularly among lower-income and caste-identified groups. Overall, outpatient and inpatient predictions from Model 2 were 8.92 (95% CI: 8.87-8.99)% and 9.59 (9.28-9.85)% higher, respectively. Nationally, this translated to an underestimation of I$390.7 (391.2-391.5) million in outpatient care and I$88.4 (86.2-90.1) million in inpatient care. The largest gaps were concentrated in the poorest and most marginalized subgroups. The SHAP analysis suggests that self-rated health (SRH), economic status (MPCE), and chronic conditions are consistently influential in predicting outpatient and inpatient visits, with some shifts in feature importance between models. ConclusionMachine learning models trained on unadjusted population data lead to algorithmic bias and risk perpetuating structural inequities by underrepresenting unmet need. Models based on fulfilled care scenarios yield more equitable and accurate projections.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The COVID-19 mortality effects of underlying health conditions in India: a modelling study 94%
- Public acceptability of non-pharmaceutical interventions to control a pandemic in the United Kingdom: a discrete choice experiment 93%
- Methodological Considerations for Linking Household and Healthcare Provider Data for Estimating Effective Coverage: A Systematic Review 93%
Similar papers in this journal
Similar papers in this journal
- Evolving Patterns of COVID-19 Mortality in US Counties: A Longitudinal Study of Healthcare, Socioeconomic, and Vaccination Associations 94%
- Listening to older voices: Results of a cross-sectional survey of older patient-reported experiences of facility-based healthcare in Nouna, Burkina Faso 93%
- All-cause excess mortality in the State of Gujarat, India, during the COVID-19 pandemic (March 2020-April 2021) 92%
Similar papers in this journal
- Progressivity of out-of-pocket costs under Australia’s universal health care system: a national linked data study 94%
- A novel methodology to measure waiting times for community-based specialist care in a public healthcare system 91%
- Long-term Public Healthcare Burden Associated with Intimate Partner Violence among Canadian Women: A Cohort Study 90%
Similar papers in this journal
- Impact of a pilot mHealth intervention on treatment outcomes of TB patients seeking care in the private sector using Propensity Scores Matching – Evidence collated from New Delhi, India 91%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 90%
- Use of Generative AI for Health Among Urban Youth in Pakistan: A Mixed-Methods Study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.