Explainable Suicide Phenotyping from Initial Psychiatric Evaluation Notes Using Reasoning Large Language Models
Li, Z.; Wang, W.; Shahani, L. R.; Selek, S.; Vieira, R. M.; Soares, J. C.; Liu, H.; Huang, M.
Show abstract
Clinical phenotyping is the process of extracting patients observable symptoms and traits to better understand their disease condition. Suicide phenotyping focuses more on behavioral and cognitive characteristics, such as suicide ideation, attempt, and self-injury, to identify suicide risks and improve interventions. In this study, we leveraged the latest reasoning models, namely 4o, o1, and o3-mini, to perform note-level multi-label classification and reasoning generation tasks using previously annotated psychiatric evaluation notes from a safety-net psychiatric inpatient hospital in Harris County, Texas. Compared with the previously finetuned GPT-3.5 model, the out-of-box reasoning models prompted with in-context learning achieved comparable and better performance, with the highest accuracy of 0.94 and F1 of 0.90. We implemented novel clinical justification generation from these models on the traditional classification tasks. This finding marked a promising direction for performing clinical phenotyping that is interpretable and actionable using smaller, efficient reasoning models.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
Similar papers in this journal
Similar papers in this journal
- Identifying Psychosis Episodes in Psychiatric Admission Notes via Rule-based Methods, Machine Learning, and Pre-Trained Language Models 94%
- Affinity Scores: An Individual-centric Fingerprinting Framework for Neuropsychiatric Disorders 92%
- Assessing psychosis risk using quantitative markers of disorganised speech 91%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 94%
- AI chatbots not yet ready for clinical use 90%
- Precision Digital Intervention for Depression Based on Social Rhythm Principles Adds Significantly to Outpatient Treatment 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.