Accurate machine learning-based CVD risk prediction in primary care may reduce the need for routine health care checks
Dziopa, K.; Eastwood, S.; Bos, D.; Kavousi, M.; Leening, M. J. G.; Beulens, J. W. J.; Harms, P. P.; Chaturvedi, N.; Asselbergs, F. W.; Schmidt, A. F.
Show abstract
BackgroundCardiovascular risk prediction models, such as PCE, QRISK3, and SCORE2 are recommended tools to guide treatment initiation/intensification in primary care. In clinical practice, the absence of one or more required predictors is common, which precludes routine application of such models. MethodsWe developed a set of partial models predicting the 10-year risk of cardiovascular disease (CVD) and major CVD (additionally considering atrial fibrillation, heart failure, and peripheral arterial disease) using combinations of 14 predictors, allowing application in settings were only a subset of variables is available. The set of partial models was evaluated across five studies jointly comprising 105,550 participants. FindingsWe trained 4,096 unique models to predict 10-year major CVD risk, observing near identical performance evaluated against CVD and major CVD. The c-statistic ranged between: quartiles (Q1) 0.71 and Q3: 0.73 across the five studies. This was comparable to the performance of the PCE (Q1: 0.70, Q3: 0.74, 10 predictors) and SCORE2 (Q1: 0.71, Q3: 0.75, 8 predictors). Due to large number of required predictors (22/23 for men/women) the QRISK3 was evaluated in a single cohort: c-statistic 0.72 (95% CI 0.72; 0.73). Model performance remained adequate when focussing on the set of partial models using 2-4 predictors: c-statistic Q1: 0.70 and Q3: 0.71. Partial models demonstrated reasonable calibration across most studies, observing a limited risk underestimation in two cohorts. Partial models excluding blood pressure and lipids demonstrated similar performance to models incorporating these variables. The set of partial models has been made available through a python-based application programming interface. InterpretationWe show that in the presence of partially missing data, clinically relevant predictions of the 10-years risk of major CVD can be obtained by using a subset of features, facilitating improved and more timely treatment decisions. FundingDutch Research Council, British Heart Foundation, UK Research and Innovation. RESEARCH IN CONTEXTO_ST_ABSEvidence before this studyC_ST_ABSBefore submitting our article on May 5, 2025, we searched PubMed articles published from database inception, using the terms "missing data" [tiab] or "incomplete data"[tiab], "cardiovascular disease" [tiab], and "risk score" [tiab] or "prediction"[tiab]. Studies unrelated to cardiovascular disease (CVD) prediction were excluded. None of the identified CVD prediction models allowed for missing input data and instead considered missing data solely at the stage of model derivation. Added value of this studyThe applicability of widely recommended cardiovascular risk prediction models, such as SCORE2 (Europe), PCE (US), and QRISK3 (UK), is constrained by the need to measure all included variables. The absence of even a single variable, such as total cholesterol used in all three aforementioned models - precludes risk prediction. For instance, among individuals aged 40 to 69 years without a history of cardiovascular disease, only 10.8% have a recorded cholesterol measurement at any point in their medical history. To overcome these limitations, this study introduces an approach using 4,096 partial models to predict 10-year risk of (major) cardiovascular disease using combinations of 14 variables, specifically designed to address the challenge of missing data. Performance was assessed across five datasets from the UK and the Netherlands. Models including between 2 - 4 predictors already provided a discriminative ability comparable to guideline-recommended models: PCE (10 predictors), SCORE2 (8 predictors), and QRISK3 (22 predictors for women, 23 for men). Implications of all the available evidenceWe show that even when only a subset of predictor variables is available, our partial models approach can make clinically relevant predictions of the 10-years risk of (major) cardiovascular disease, enabling earlier and more effective treatment decisions. The set of partial models are accessible through a python-based API, allowing for integration in personal or clinical care dashboards.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of a risk prediction algorithm for high-risk populations combining genetic and conventional risk factors of cardiovascular disease 95%
- AORTA Gene: Polygenic prediction improves detection of thoracic aortic aneurysm 94%
- Lipoprotein(a) and cardiovascular disease: prediction, attributable risk fraction and estimating benefits from novel interventions 93%
Similar papers in this journal
Similar papers in this journal
- Nationwide prediction of type 2 diabetes comorbidities 94%
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 93%
- A multi-omics study of circulating phospholipid markers of blood pressure 92%
Similar papers in this journal
- Heritability of cardiovascular health across three generations in South Africa: the Birth to Twenty-Plus cohort 94%
- History of coronary heart disease increases the mortality rate of COVID-19 patients: a nested case-control study 93%
- Body mass index and all-cause mortality in HUNT and UK Biobank studies: revised non-linear Mendelian randomization analyses 92%
Similar papers in this journal
- Mendelian randomisation for mediation analysis: current methods and challenges for implementation 93%
- A comparison of regression discontinuity and propensity score matching to estimate the causal effects of statins using electronic health records 93%
- Associations of polygenic inheritance of physical activity with aerobic fitness, cardiometabolic risk factors and diseases: the HUNT Study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.