Evaluating the effect of mental health fine-tuning relative to other model characteristics on LLM safety performance
Kalinich, M.; Luccarelli, J.; Santa Maria, J.; Williams, G.; Moss, F.; Torous, J.
Show abstract
Large language models (LLMs) are increasingly used in mental health applications, yet it remains unclear whether mental health-specific fine-tuning meaningfully improves safety-relevant performance beyond gains from model scale or architecture. We evaluated 127 publicly available open-source LLMs across three model families, multiple architecture generations, parameter scales (270M-70B), and fine-tuning strategies on three psychiatrist-reviewed synthetic classification tasks: suicidal ideation detection, identification of user requests for therapy, and detection of explicit therapy-like interactions in multi-turn conversations. Performance was summarized using F1 score, with multivariable regression and paired comparisons used to estimate independent effects of model characteristics. Across tasks, newer architectures and larger models consistently showed superior performance. General instruction tuning improved detection of therapy requests and engagement, whereas mental health-specific, medical, or safety fine-tuning conferred no consistent benefit and were sometimes associated with reduced performance. These findings suggest that baseline model capability is more consequential than domain-specific fine-tuning for certain safety-relevant mental health classification tasks, underscoring the importance of careful model selection and task-specific evaluation.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 92%
- Improving ascertainment of suicidal ideation and suicide attempt with natural language processing 92%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.