Medication-Stratified Analysis of LDL-C Equation Miscalibration in Diabetes: Evidence from the All of Us Research Program and a Medication-Agnostic Machine-Learning Correction
Doku, R.; Osafo, N. Y.; Kwagyan, J.
Show abstract
ObjectiveStandard LDL-C equations were derived in cohorts largely untreated with modern combination diabetes therapies. With medication-treated patients comprising 84% on statins, 53% on insulin, and 25% on GLP-1 receptor agonists--often in combination--we quantified medication-specific miscalibration in LDL-C equations and evaluated a machine learning correction that operates without requiring medication data. Research Design and MethodsUsing All of Us Research Program data (n=3,477; test =696), we compared Friedewald, Martin-Hopkins, and Sampson (NIH) Equation 2 against direct LDL-C measurements. We developed a stacked ensemble model (elastic net, random forest, XGBoost, neural network) trained solely on routine laboratory values. Accuracy was assessed within medication groups allowing for combination therapy: insulin users, GLP-1 users, and statin users. Primary endpoints: mean absolute error (MAE) with 95% bootstrap confidence intervals and calibration (ordinary least squares regression of true on predicted LDL-C). Secondary endpoint: Net Reclassification Index at 100 mg/dL. ResultsAmong 696 test participants, 587 (84%) used statins, 366 (53%) insulin, and 175 (25%) GLP-1 agonists. Patients on triple therapy (insulin+GLP-1+statin) showed the most severe miscalibration: Friedewald slope 0.29, representing 71% compression of the prediction range. In all GLP-1 users (77% also on insulin), standard equations severely underestimated LDL-C with calibration slopes of 0.42-0.48 versus ideal 1.0. Specifically, Friedewald showed slope 0.42 (95% CI 0.27-0.56) with intercept +62 mg/dL; Sampson (NIH) Equation 2 slope 0.48 (0.32-0.64) with intercept +55 mg/dL; Martin-Hopkins slope 0.47 (0.31-0.63) with intercept +55 mg/dL. The machine learning model maintained better calibration (slope 0.83 [0.56-1.09]; intercept -2.2 mg/dL) and reduced MAE by 17% versus Friedewald. Insulin users showed similar improvement: Friedewald slope 0.55 (0.45-0.65) versus the machine learning (ML) model 0.95 (0.78-1.12), with 16% lower error. The medication-by-triglyceride interaction was significant (p=0.002). In patients with insulin exposure and triglycerides [≥]200 mg/dL, Net Reclassification Index was 0.240 versus 0.022 overall, indicating greater misclassification risk in hypertriglyceridemia. ConclusionsStandard LDL-C equations systematically underestimate true levels in medication-treated diabetes patients, with errors greatest in combination therapy. A machine learning model trained on routine laboratories--without medication data--achieved near-ideal calibration (slopes 0.83-1.03) and reduced errors by 8-20% across medication groups. These observational findings suggest direct LDL-C measurement or ML-assisted correction should be considered when equation estimates approach treatment thresholds, particularly for patients on combination therapy.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Heterogeneity of Treatment Effects Across Nine Glucose-Lowering Drug Classes in Type 2 Diabetes: Extension of the LEGEND-T2DM Network Study 93%
- Glucagon-like peptide-1 receptor agonists modestly reduced blood pressure among patients with and without diabetes mellitus: A meta-analysis and meta-regression 92%
- Comparative Effects of Weight Loss and Incretin-Based Therapies on Endothelial Vasodilatory and Fibrinolytic Function 92%
Similar papers in this journal
Similar papers in this journal
- Cardiovascular risk prediction in type 2 diabetes: a comparison of 22 risk scores in primary care setting 93%
- Precision medicine in Type 2 Diabetes: Targeting SGLT2-inhibitor Treatment For Kidney Protection 92%
- Phenotype-based targeted treatment of SGLT2 inhibitors and GLP-1 receptor agonists in type 2 diabetes 92%
Similar papers in this journal
- Differential impact of Covid-19 on incidence of diabetes mellitus and cardiovascular diseases in acute, post-acute and long Covid-19: population-based cohort study in the United Kingdom 91%
- Harnessing the power of polygenic risk scores to predict type 2 diabetes and its subtypes in a high-risk population of British Pakistanis and Bangladeshis in a routine healthcare setting 90%
- Predicting and elucidating the etiology of fatty liver disease using a machine learning-based approach: an IMI DIRECT study 89%
Similar papers in this journal
- Statin treatment effectiveness and the SLCO1B1 *5 reduced function genotype: long-term outcomes in women and men 91%
- Calcium Channel Blockers: clinical outcome associations with reported pharmacogenetics variants in 32,000 patients 91%
- Antihypertensive Medications and COVID-19 Diagnosis and Mortality: Population-based Case-Control Analysis in the United Kingdom 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.