Back

A Machine Learning Approach to Predicting Heterogeneous Lipid Responses to Statins in Type 2 Diabetes

Harris, C. W.; Olshvang, D.; Chellappa, R.; Santhanam, P.

2026-01-02 endocrinology
10.64898/2026.01.01.26343317 medRxiv
Show abstract

BackgroundInter-individual variability in lipid response to statin therapy poses a major challenge in cardiovascular risk reduction, particularly among patients with type 2 diabetes mellitus (T2DM), who exhibit complex dyslipidemia and elevated cardiovascular risk. While randomized trials provide population-level estimates of treatment efficacy, individualized prediction of lipid changes, and corresponding uncertainty, remains limited. Using rich clinical data from the Action to Control Cardiovascular Risk in Diabetes (ACCORD) study, this work sought to quantify and predict person-level changes in multiple lipid fractions following statin initiation. MethodsAmong 3,509 ACCORD participants with T2DM who were newly initiated on simvastatin and had 12-month lipid follow-up, we developed machine learning models to predict changes in LDL-C, HDL-C, triglycerides, VLDL, and total cholesterol using only baseline clinical features. Five model classes were evaluated using nested cross-validation; Lasso regression emerged as the best-performing model for LDL-C. To provide clinically interpretable and statistically valid uncertainty estimates, we applied distribution-free conformal prediction to generate individual-level prediction intervals at multiple confidence levels. Feature importance was quantified using SHAP values. ResultsLongitudinal lipid responses showed substantial heterogeneity across individuals. Predictability varied markedly by lipid fraction: LDL-C (2 = 0.52; = 18.3 /) and total cholesterol (2 = 0.41) were moderately estimable from baseline data, whereas HDL-C (2{approx} 0.04) exhibited minimal predictability, and triglycerides and VLDL demonstrated intermediate but volatile behavior. Baseline lipid values were the strongest predictors across all outcomes, followed by markers of statin exposure and metabolic status. Conformal prediction achieved accurate empirical coverage across all lipid fractions ({approx} 90% at nominal 90% confidence), with interval widths scaling appropriately to biological variability. Undercoverage occurred primarily among individuals predicted to have minimal lipid changes. ConclusionsMachine learning models can meaningfully predict LDL-C and total cholesterol responses following statin initiation in patients with T2DM, while HDL-C and triglyceride responses remain dominated by unmeasured or stochastic factors. Conformal prediction adds a crucial layer of reliability by providing calibrated, patient-specific uncertainty intervals, supporting safer and more transparent clinical decision-making. These findings highlight both the promise and the current limitations of EHR-based precision lipid management and underscore the need to incorporate genetic, metabolic, and behavioral data to capture the full spectrum of lipid response variability.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.