Predicting Intentional Self-Harm Following Psychiatric Discharge in Catalonia, Spain: Machine Learning Models from Linked Registry Data
Alayo, I.; Pujol, O.; Amigo, F.; Ballester, L.; Cirici Amell, R.; Contaldo, S. F.; Ferrer, M.; Guinart, D.; Latorre, L.; Leis, A.; Lopez Fernandez, M.; Mayer, M. A.; Pastor, M.; Pena-Salazar, C.; Portillo-Van Diest, A.; Ramirez-Anguita, J. M.; Sanz, F.; Alonso, J.; Kessler, R. C.; Mehlum, L.; Palao, D.; Perez Sola, V.; Vilagut, G.; Mortier, P.
Show abstract
IntroductionPatients recently discharged from psychiatric hospitalization are at increased risk of intentional self-harm, including suicide. Using linked population-based registry data from Catalonia, Spain, we developed machine learning-based prediction models for post-discharge intentional self-harm across different follow-up horizons, sex, and age groups, and evaluated their generalizability and robustness with multiple validation strategies. MethodsRetrospective cohort study including 41,827 individuals accounting for 71,865 psychiatric hospitalizations with discharge at age [≥]10 years, between January 1, 2015, and December 31, 2018, in Catalonia, Spain, with follow-up until December 31, 2019. Primary outcome was intentional self-harm (fatal or non-fatal) within 7, 30, 90, 180, and 365 days post-discharge. Models incorporated 247 predictors from electronic health records, including sociodemographic characteristics, mental and physical disorder categories, categories of dispensed psychotropic medication, and history of self-harm and psychiatric hospitalization. Model performance was evaluated using the area under the receiver operating characteristic curve (AUCROC) and the area under the precision-recall curve (AUCPR). Predictor importance was assessed using Shapley Additive Explanations (SHAP). ResultsWithin 365 days, 4,901 hospitalizations (6.8%) were followed by intentional self-harm. The 365-day model trained on the full cohort achieved a AUCROC of 0.819, in the test sample with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence. This model generalized well across event horizons and sex-age strata, outperforming subgroup-specific models when data sparsity limited performance. Separate models trained by event horizons, and stratified by sex, and sex-age groups achieved a median AUCROC of 0.775 (IQR 0.764-0.808), with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence (IQR 4.5-6.2). Key predictors included the recency of the last registered diagnosis of depressive episodes, recurrent depression, adjustment disorders, and schizophrenia, as well as recent SSRI dispensation and the number of childhood-onset disorder and musculoskeletal disease diagnoses in the previous five years. Predictor importance varied considerably across sex-age strata, with smaller differences across horizons. Subject-level and temporal split validation strategies reduced performance (AUCROC 0.711-0.746), though estimates remained clinically informative (2.8-3.1-fold improvement over baseline prevalence). ConclusionsMachine learning models using routinely collected health records predicted intentional self-harm after psychiatric hospitalization with good discrimination and clinically meaningful precision-recall performance. A single 365-day model generalized well across horizons and demographic groups, suggesting that one broadly trained model may provide a pragmatic and scalable approach for clinical implementation.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predictors of suicidal thoughts and behavior in children: results from penalized logistic regression analyses in the ABCD study 95%
- Analysis of diagnosis instability in electronic health records reveals diverse disease trajectories of severe mental illness 95%
- Associations between antipsychotic use, substance use and relapse risks in patients with schizophrenia - real-world evidence from two national cohorts 94%
Similar papers in this journal
- Predicting remission after internet-delivered psychotherapy in patients with depression using machine learning and multi-modal data 95%
- Association between Psychotropic Medications Functionally Inhibiting Acid Sphingomyelinase and reduced risk of Intubation or Death among Individuals with Mental Disorder and Severe COVID-19: an Observational Study 95%
- Prognostic models predicting transition to psychotic disorder using blood-based biomarkers: a systematic review and critical appraisal 94%
Similar papers in this journal
- Identifying Potential Causal Risk Factors for Self-Harm: A Polygenic Risk Scoring and Mendelian Randomisation Approach 94%
- Clinical benefits of modifying the evening light environment in an acute psychiatric unit: A single-centre, two-arm, parallel-group, pragmatic effectiveness randomised controlled trial 94%
- Risk of common psychiatric disorders, suicidal behaviours and premature mortality following violent victimisation: A matched cohort and sibling-comparison study of 127,628 people who experienced violence in Finland and Sweden 94%
Similar papers in this journal
- Longitudinal evolution of the transdiagnostic prodrome to severe mental disorders: a dynamic temporal network analysis informed by natural language processing and electronic health records 94%
- Cross-phenotype relationship between opioid use disorder and suicide attempts: new evidence from polygenic association and Mendelian randomization analyses 93%
- Correlates of suicidal behaviors and genetic risk among United States veterans with schizophrenia or bipolar I disorder 93%
Similar papers in this journal
- Life-years lost associated with mental illness: a cohort study of beneficiaries of a South African medical insurance scheme 94%
- Multidimensional apathy: A simple and inclusive clinical marker of youth mental health—A longitudinal study 94%
- Depression is Associated with Treatment Response Trajectories in Adults with Prolonged Grief Disorder: A Machine Learning Analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.