A machine learning model to aid detection of familial hypercholesterolaemia
Gratton, J.; Futema, M.; Humphries, S. E.; Hingorani, A. D.; Finan, C.; Schmidt, A. F.
Show abstract
2.TEXT ABSTRACT AND KEYWORDSO_ST_ABSBackground and AimsC_ST_ABSPeople with monogenic familial hypercholesterolaemia (FH) are at an increased risk of premature coronary heart disease and death. Currently there is no population screening strategy for FH, and most carriers are identified late in life, delaying timely and cost-effective interventions. The aim was to derive an algorithm to improve detection of people with monogenic FH. MethodsA penalised (LASSO) logistic regression model was used to identify predictors that most accurately identified people with a higher probability of FH in 139,779 unrelated participants of the UK Biobank, including 488 FH carriers. Candidate predictors included information on medical and family history, anthropometric measures, blood biomarkers, and an LDL-C polygenic score (PGS). Model derivation and evaluation was performed using a random split of 80% training and 20% testing data. ResultsA 14-variable algorithm for FH was derived, where the top five variables included triglyceride, LDL-C, and apolipoprotein A1 concentrations, self-reported statin use, and an LDL-C PGS. Model evaluation in the test data resulted in an area under the curve (AUC) of 0.77 (95% CI: 0.71; 0.83), and appropriate calibration (calibration-in-the-large: -0.07 (95% CI: -0.28; 0.13); calibration slope: 1.02 (95% CI: 0.85; 1.19)). Employing this model to prioritise people with suspected monogenic FH is anticipated to reduce the number of people requiring sequencing by 88% compared to a population-wide sequencing screen, and by 18% compared to prioritisation based on LDL-C and statin use. ConclusionsThe detection of individuals with monogenic FH can be improved with the inclusion of additional non-genetic variables and a PGS for LDL-C.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Coronary Artery Disease Risk of Familial Hypercholesterolemia Genetic Variants Independent of Historical Cholesterol Exposure 96%
- Population genomic screening leads to improved lipid management in patients with familial hypercholesterolemia 95%
- Validation of genome-wide polygenic risk scores for coronary artery disease in French Canadians 95%
Similar papers in this journal
- Adverse pregnancy outcomes and coronary artery disease risk: A negative control Mendelian randomization study 93%
- Genetically downregulated interleukin-6 signaling is associated with a favorable cardiometabolic profile: a phenome-wide association study 92%
- Genetically Predicted IL-18 Inhibition and Risk of Cardiovascular Events: A Mendelian Randomization Study 92%
Similar papers in this journal
- Polygenic Hyperlipidemias and Coronary Artery Disease Risk 96%
- LDLR Variant Classification for Improved Cardiovascular Risk Prediction in Familial Hypercholesterolemia 95%
- Apolipoprotein E genotype, lifestyle and coronary artery disease: gene-environment interaction analyses in the UK Biobank population 93%
Similar papers in this journal
- Educational attainment as a modifier of the effect of polygenic scores for cardiovascular risk factors: cross-sectional and prospective analysis of UK Biobank 95%
- High-throughput multivariable Mendelian randomization analysis prioritizes apolipoprotein B as key lipid risk factor for coronary artery disease 95%
- Genetic Risk Scores for Cardiometabolic Traits in Sub-Saharan African Populations 94%
Similar papers in this journal
- Genetically-proxied therapeutic inhibition of antihypertensive drug targets and risk of common cancers 93%
- Haplotype genetic score analysis in 10,734 mother/infant pairs reveals complex maternal and fetal genetic effects underlying the associations between maternal phenotypes, birth outcomes and adult phenotypes 91%
- Assessing a causal relationship between circulating lipids and breast cancer risk: Mendelian randomization study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.