Back

Development and validation of gradient boosting decision tree models for predicting care needs using a long-term care database in Japan

Lin, H.-R.; Fujiwara, K.; Minoru, S.; Ishiyama, K.; Ikeda-Sonoda, S.; Takahashi, A.; Miyata, H.

2021-01-26 health informatics
10.1101/2021.01.20.21250146 medRxiv
Show abstract

ObjectiveThe purpose of the study was to develop machine learning models using data from long-term care (LTC) insurance claims and care needs certifications to predict the individualized future care needs of each older adult. MethodsWe collected LTC insurance-related data in the form of claims and care needs certification surveys from a municipality of Kanagawa Prefecture from 2009 to 2018. We used care needs certificate applications for model generation and the validation sample to build gradient boosting decision tree (GBDT) models to classify if 1) the insureds care needs either remained stable or decreased or 2) the insureds care needs increased after three years. The predictive model was trained and evaluated via k-fold cross-validation. The performance of the predictive model was observed in its accuracy, precision, recall, F1 score, area under the receiver-operator curve, and confusion matrix. ResultsLong-term care certificate applications and claim data from 2009-2018 were associated with 92,239 insureds with a mean age of 86.1 years old at the time of application, of whom 67% were female. The classifications of increase in care needs after three years were predicted with AUC of 0.80. ConclusionsMachine learning is a valuable tool for predicting care needs increases in Japans LTC insurance system, which can be used to develop more targeted and efficient interventions to proactively reduce or prevent further functional deterioration, thereby helping older adults maintain a better quality of life.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.