Back

Suicide Death Predictive Models using Electronic Health Record Data

Srikanth, S.; Montoya, L.; Turnure, M. M.; Pence, B. W.; Fulcher, N.; Gaynes, B. N.; Goldston, D.; Carey, T.; Ranapurwala, S. I.

2024-09-27 epidemiology
10.1101/2024.09.26.24314402 medRxiv
Show abstract

In the realm of medical research, particularly in the study of suicide risk assessment, the integration of machine learning techniques with traditional statistics methods has become increasingly prevalent. This paper used data from the UNC EHR system from 2006 to 2020 to build models to predict suicide-related death. The dataset, with 1021 cases and 10185 controls consisted of demographic variables and short-term informa-tion, on the subjects prior diagnosis and healthcare utilization. We examined the efficacy of the super learner ensemble method in predicting suicide-related death lever-aging its capability to combine multiple predictive algorithms without the necessity of pre-selecting a single model. The study compared the performance of the super learner against five base models, demonstrating its superiority in terms of cross-validated neg-ative log-likelihood scores. The super learner improved upon the best algorithm by 60% and the worst algorithm by 97.5%. We also compared the cross-validated AUCs of the models optimized to have the best AUC to highlight the importance of the choice of risk function. The results highlight the potential of the super learner in complex predictive tasks in medical research, although considerations of computational expense and model complexity must be carefully managed.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.