Algorithmic Fairness of QPrediction Cardiometabolic Risk Prediction Models
Huynh, I.; Utochkin, D.; Katsiferis, A.; Nguyen, T.-L.; Varga, T. V.
Show abstract
BackgroundThere is a lack of fairness evaluation of clinical prediction models used in routine practice. This study aimed to evaluate the fairness of three established risk prediction models belonging to the QPrediction family: QRISK3, QDiabetes, and QStroke, which estimate 10-year risks for cardiovascular disease, type 2 diabetes, and stroke, respectively. Methods and findingsWe used data from the UK Biobank, a large-scale prospective cohort study comprising 502,131 participants. To assess calibration, the predicted risks were compared to the observed cumulative incidence rates, and survival Brier scores were calculated as a summary metric of model calibration. To assess discrimination, we calculated overall and group-specific true positive, true negative, false positive, and false negative rates (TPR, TNR, FPR, FNR). We compared the rates between subgroups of demographics and considered prediction parity when values were close to each other to a prespecified degree. Overall, all models showed systematic overprediction of risks for the UK Biobank validation cohorts, with the degree of overprediction varying between the models. Prediction parity was not reached in FNR and TNR in terms of ethnicity, education level, average household income before taxation, and Townsend deprivation score. For immigration status, there were no major differences in the discrimination metrics, and prediction parity was achieved. For QDiabetes, the trends were similar, apart from ethnicity, where the opposite was observed as for QRISK3. ConclusionsInclusion of ethnicity and deprivation in the models may have contributed to more consistent calibration across subgroups. Variables were measured only at baseline, while many social and economic conditions can shift over time. Although including social determinants of health can improve predictive performance, the models are typically deployed in a clinical setting to guide individual-level decision making, which can inadvertently shift responsibility for health outcomes onto patients.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 96%
- Latent class regression improves the predictive acuity and clinical utility of survival prognostication amongst chronic heart failure patients. 95%
- Patterns of rates of mortality in the Clinical Practice Research Datalink 93%
Similar papers in this journal
- Sociodemographic Characteristics and Longitudinal Progression of Multimorbidity: A Multistate Modelling Analysis of a Large Primary Care Records Dataset in England 95%
- Factors associated with excess all-cause mortality in the first wave of COVID-19 pandemic in the UK: a time-series analysis using the Clinical Practice Research Datalink 94%
- Rapid Epidemiological Analysis of Comorbidities and Treatments as risk factors for COVID-19 in Scotland (REACT-SCOT): a population-based case-control study 92%
Similar papers in this journal
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 94%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 92%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 92%
Similar papers in this journal
- Polygenic Risk Score Improves the Accuracy of a Clinical Risk Score for Coronary Artery Disease 93%
- Ethnic differences in COVID-19 mortality in the second and third waves of the pandemic in England during the vaccine roll-out: a retrospective, population-based cohort study 92%
- Coding Long COVID: Characterizing a new disease through an ICD-10 lens 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.