Back

Developing a Health Score and Predicting Disease Risks Using DKAbio-Clusters

Cheng, K.-F.; Yang, Y.-H.; Su, C.-H.; Tsai, M.-C.

2024-06-18 health informatics
10.1101/2024.06.16.24308995 medRxiv
Show abstract

In our research, accurately estimating the morbidity of individuals with specific conditions, plays a pivotal role in enhancing healthcare delivery systems. Introducing DKABio-clusters, we delve into their distinct characteristics, showcasing their profound implications for healthcare management. A primary focus of DKABio-clusters lies in developing a unique health assessment tool, termed DKABio-HS, alongside predictive risk analysis. DKABio-HS facilitates the computation of a comprehensive "disease-related" score, condensing an individuals health status into a singular numerical value. Our investigation reveals the remarkable consistency of this health score, with minimal variations observed between training and validation datasets (mean absolute percentage errors within 0 to 10 years remaining below 0.1%, with all mean absolute percentage errors ranging between 1.2-1.6%). A higher health score denotes better health or reduced disease risk, diminishing with age or the presence of multiple diseases. Utilizing this health score, we establish a classification framework termed the "disease map," enabling precise differentiation of individuals across various health states. Through this framework, individuals without diseases can be categorized as either healthy or sub-healthy, facilitating tailored health management strategies for preventive interventions. Our analysis indicates that individuals classified as sub-healthy exhibit significantly elevated disease risks compared to those deemed healthy (Female (male) 5-year risks of developing at least one disease are 29% vs. 15% (29% vs. 16.5%)). Furthermore, leveraging a carefully selected set of health variables, we can delineate the distribution of DKABio-clusters and concurrently predict the 10-year risks associated with 15 diseases/conditions. Validating the predictive capabilities of our model, we compare predicted risks with true risks derived from extensive datasets, demonstrating non-statistically significant differences in the majority of cases. All analyses are grounded in data sourced from the National Health Insurance Research Database (more than 2 million participants) released by the National Health Research Institute, Taiwan and the Mei Jau Health Management Institution database (more than 0.75 million participants), spanning the years 2000 to 2016 in Taiwan.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.