Back

Dynamic Stroke Risk Stratification via Machine Learning: A Multi-Level Single-Center Study

Liao, Z.; Yan, Z.; Huang, J.; Chen, X.; Yang, Y.

2025-12-05 public and global health
10.64898/2025.12.03.25341598 medRxiv
Show abstract

BackgroundStroke is a leading global public health challenge and the second leading cause of death worldwide. In China, its burden continues to escalate amid population aging and a growing prevalence of unhealthy lifestyles. Traditional static stroke risk prediction models, constrained by cross-sectional data, fail to capture dynamic changes in physiological parameters and behavioral factors, resulting in inherent limitations. This study therefore aimed to construct and validate a multi-tiered dynamic stroke risk prediction system. MethodsA single-center longitudinal prospective cohort study (2018-2022) enrolled community-dwelling populations and outpatients who completed three consecutive follow-ups. Three sequential machine learning models were developed, targeting general population screening, high-risk population refinement, and longitudinal population monitoring, respectively. Stratified cross-validation was used for model validation (10-fold for Model 1, 5-fold for Models 2 and 3), with the area under the curve (AUC) as the primary evaluation metric. ResultsModel 1 (for general population high-risk conversion prediction) achieved an AUC of 0.835, which significantly reduced the missed detection rate of individuals with "normal static indicators but abnormal dynamic trends." Model 2 (for high-risk population stroke onset prediction) had LGB_Conservative as its optimal algorithm, with an AUC of 0.8479. Model 3 (for longitudinal population dual-outcome prediction) showed an AUC of 0.761, with stroke event rates of 3.1% in the low-risk group and 95.2% in the very high-risk group. ConclusionThis multi-tiered dynamic prediction system effectively addresses the limitations of traditional static models, yet requires external validation using multi-center data to confirm its generalizability. It provides a novel tool for personalized stroke prevention in clinical and public health practice.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.