Back

Development of a future prediabetes risk assessment model for individuals with normal glucose levels using efficiency scores obtained from data envelopment analysis

Nakamura, S.; Inoue, R.; Narimatsu, H.

2025-05-21 epidemiology
10.1101/2025.05.21.25328057 medRxiv
Show abstract

BackgroundIdentifying healthy individuals at risk of prediabetes for primary prevention is crucial, as current tools often focus on secondary prevention. We investigated whether efficiency scores, derived from data envelopment analysis (DEA), predict prediabetes development in a healthy population. MethodsThis historical cohort study analyzed annual health checkup data. Cox proportional hazards analysis assessed the relationship between efficiency scores and incident prediabetes. A classification tree analysis was also performed, incorporating efficiency scores, hemoglobin A1c (HbA1c), and other diabetes-related variables. ResultsThe cohort comprised 923 individuals (49.7% female), with a mean efficiency score of 0.72 (0.07). During follow-up, 175 participants developed prediabetes (79.3 per 1,000 person-years). A 0.1-point increase in efficiency score was associated with an adjusted hazard ratio of 0.51 (95% CI 0.39-0.68, p < 0.0001) for prediabetes, while a 0.1% increase in HbA1c yielded an adjusted hazard ratio of 2.26 (95% CI 1.88-2.71, p < 0.0001). The classification tree identified a high-risk group of 31 individuals (3.4%) with 12.1% sensitivity and 98.7% specificity. DiscussionEfficiency scores are linked to the 3-year risk of prediabetes in healthy subjects. The combined use of DEA and classification tree analysis presents a potentially valuable approach for primary prevention strategies in clinical practice.

Published in PNAS Nexus (predicted rank #17) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.