Back

Algorithmic Fairness of QPrediction Cardiometabolic Risk Prediction Models

Huynh, I.; Utochkin, D.; Katsiferis, A.; Nguyen, T.-L.; Varga, T. V.

2025-08-24 public and global health
10.1101/2025.08.20.25333669 medRxiv
Show abstract

BackgroundThere is a lack of fairness evaluation of clinical prediction models used in routine practice. This study aimed to evaluate the fairness of three established risk prediction models belonging to the QPrediction family: QRISK3, QDiabetes, and QStroke, which estimate 10-year risks for cardiovascular disease, type 2 diabetes, and stroke, respectively. Methods and findingsWe used data from the UK Biobank, a large-scale prospective cohort study comprising 502,131 participants. To assess calibration, the predicted risks were compared to the observed cumulative incidence rates, and survival Brier scores were calculated as a summary metric of model calibration. To assess discrimination, we calculated overall and group-specific true positive, true negative, false positive, and false negative rates (TPR, TNR, FPR, FNR). We compared the rates between subgroups of demographics and considered prediction parity when values were close to each other to a prespecified degree. Overall, all models showed systematic overprediction of risks for the UK Biobank validation cohorts, with the degree of overprediction varying between the models. Prediction parity was not reached in FNR and TNR in terms of ethnicity, education level, average household income before taxation, and Townsend deprivation score. For immigration status, there were no major differences in the discrimination metrics, and prediction parity was achieved. For QDiabetes, the trends were similar, apart from ethnicity, where the opposite was observed as for QRISK3. ConclusionsInclusion of ethnicity and deprivation in the models may have contributed to more consistent calibration across subgroups. Variables were measured only at baseline, while many social and economic conditions can shift over time. Although including social determinants of health can improve predictive performance, the models are typically deployed in a clinical setting to guide individual-level decision making, which can inadvertently shift responsibility for health outcomes onto patients.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.