Privacy-preserving local language models accurately identify the presence and timing of self-harm in electronic mental health records
Kormilitzin, A.; Joyce, D. W.; Tsiachristas, A.; Borschmann, R.; Kapur, N.; Geulayov, G.
Show abstract
BackgroundSelf-harm, defined as intentional self-injury or self-poisoning irrespective of motivation, is the strongest risk factor for suicide and an important outcome for mental health care. Although highly prevalent in clinical populations, it is often imprecisely captured in routinely collected clinical data, where it is typically recorded and stored in unstructured, free-text format. Contemporary language models, such as GPT (OpenAI), Gemini (Google) and Claude (Anthropic) can analyse free-text clinical notes, but such cloud-based commercial and closed-source models may violate data governance of processing sensitive patient data. ObjectiveWe evaluated whether a privacy-preserving language model running entirely within an institutions secure computing infrastructure (here, the UK National Health Service) could accurately identify the presence and timing of self-harm using electronic health records (EHRs) from secondary mental healthcare. MethodsA random sample of 1,352 clinical notes from the EHRs of people with a confirmed diagnosis of a psychiatric disorder, according to the International Classification of Diseases (ICD-10), was selected from Oxford Health NHS Foundation Trust. Each clinical note was annotated for (i) presence or absence of self-harm and (ii) its recency ([≤]90 days vs. >90 days vs. unknown), constituting the gold-standard data set for model development and validation. Privacy-preserving locally served 27-billion-parameter Gemma 3 language model ( Gemma3-27b) was used as a core language model to identify self-harm along with its timing to generate a structured output per clinical record. The performance of Gemma3-27b model was compared against a strong baseline multi-label text classification model based on RoBERTa (Robustly Optimized BERT Pretraining Approach, a transformer-based language model) architecture. Model performance was evaluated using precision, recall, and the F1-score (harmonic mean of precision and recall), with 95% confidence intervals estimated from 1,000 bootstrap samples with replacement. ResultsThe Gemma3-27b model outperformed the RoBERTa classifier across all categories. Gemma3-27b achieved Precision = 0.92, Recall = 0.92 (sensitivity), and F1-score of 0.92 for notes containing self-harm, and Precision = 0.97, Recall = 0.97 (specificity), and F1-score of 0.97 for notes without self-harm. The performance of Gemma3-27b for recent self-harm was Precision = 0.84, Recall = 0.75, and F1-score of 0.79. The global weighted F1-score of Gemma3-27b across all categories was 0.88, compared to 0.85 for RoBERTa. ConclusionsThe prompt-engineered local language model outperformed the multi-label text classification RoBERTa model on the task of ascertaining self-harm events and their timing, while requiring significantly fewer high-fidelity annotated data. This approach offers a practical solution for improving self-harm case ascertainment whilst adhering to strict data governance rules to potentially transform clinical monitoring and response to this critical public health challenge.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 93%
- Using indication embeddings to represent patient health for drug safety studies 92%
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 92%
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 91%
Similar papers in this journal
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 92%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 91%
- Can we trust the prediction model? Demonstrating the importance of external validation by investigating the COVID-19 Vulnerability (C-19) Index across an international network of observational healthcare datasets 91%
Similar papers in this journal
- Closing the accessibility gap to mental health treatment with a conversational AI-enabled self-referral tool 91%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 91%
- Fostering transparent medical image AI via an image-text foundation model grounded in medical literature 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.