Back

Acta Psychiatrica Scandinavica

Wiley

All preprints, ranked by how well they match Acta Psychiatrica Scandinavica's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Suicide prediction with natural language processing of electronic health records

Korda, A.

2023-09-29 psychiatry and clinical psychology 10.1101/2023.09.28.23296268 medRxiv
Top 0.1%
9.5%
Show abstract

Suicide attempts are one of the most challenging psychiatric outcomes and have great importance in clinical practice. However, they remain difficult to detect in a standardised way to assist prevention because assessment is mostly qualitative and often subjective. As digital documentation is increasingly used in the medical field, Electronic Health Records (EHRs) have become a source of information that can be used for prevention purposes, containing codified data, structured data, and unstructured free text. This study aims to provide a quantitative approach to suicidality detection using EHRs, employing natural language processing techniques in combination with deep learning artificial intelligence methods to create an algorithm intended for use with medical documentation in German. Using psychiatric medical files from in-patient psychiatric hospitalisations between 2013 and 2021, free text reports will be transformed into structured embeddings using a German trained adaptation of Word2Vec, followed by a Long-Short Term Memory (LSTM) - Convolutional Neural Network (CNN) approach on sentences of interest. Text outside the sentences of interest will be analysed as context using a fixed size ordinally-forgetting encoding (FOFE) before combining these findings with the LSTM-CNN results in order to label suicide related content. This study will offer promising ways for automated early detection of suicide attempts and therefore holds opportunities for mental health care.

2
Pathway and network embedding methods for prioritizing psychiatric drugs

Pershad, Y.; Guo, M.; Altman, R. B.

2019-08-13 bioinformatics 10.1101/728055 medRxiv
Top 0.1%
8.3%
Show abstract

One in five Americans experience mental illness, and roughly 75% of psychiatric prescriptions do not successfully treat the patients condition. Extensive evidence implicates genetic factors and signaling disruption in the pathophysiology of these diseases. Changes in transcription often underlie this molecular pathway dysregulation; individual patient transcriptional data can improve the efficacy of diagnosis and treatment. Recent large-scale genomic studies have uncovered shared genetic modules across multiple psychiatric disorders--providing an opportunity for an integrated multi-disease approach for diagnosis. Moreover, network-based models informed by gene expression can represent pathological biological mechanisms and suggest new genes for diagnosis and treatment. Here, we use patient gene expression data from multiple studies to classify psychiatric diseases, integrate knowledge from expert-curated databases and publicly available experimental data to create augmented disease-specific gene sets, and use these to recommend disease-relevant drugs. From Gene Expression Omnibus, we extract expression data from 145 cases of schizophrenia, 82 cases of bipolar disorder, 190 cases of major depressive disorder, and 307 shared controls. We use pathway-based approaches to predict psychiatric disease diagnosis with a random forest model (78% accuracy) and derive important features to augment available drug and disease signatures. Using protein-protein-interaction networks and embedding-based methods, we build a pipeline to prioritize treatments for psychiatric diseases that achieves a 3.4-fold improvement over a background model. Thus, we demonstrate that gene-expression-derived pathway features can diagnose psychiatric diseases and that molecular insights derived from this classification task can inform treatment prioritization for psychiatric diseases.

3
Multimodal Data Integration Advances Longitudinal Prediction of the Naturalistic Course of Depression and Reveals a Multimodal Signature of Disease Chronicity

Habets, P. C.; Thomas, R. M.; Milaneschi, Y.; Jansen, R.; Pool, R.; Peyrot, W. J.; Penninx, B. W.; Meijer, O. C.; van Wingen, G. A.; Vinkers, C. H.

2023-01-11 genomics 10.1101/2023.01.10.523383 medRxiv
Top 0.1%
8.0%
Show abstract

The ability to individually predict disease course of major depressive disorder (MDD) is essential for optimal treatment planning. Here, we use a data-driven machine learning approach to assess the predictive value of different sets of biological data (whole-blood proteomics, lipid-metabolomics, transcriptomics, genetics), both separately and added to clinical baseline variables, for the longitudinal prediction of 2-year MDD chronicity (defined as presence of MDD diagnosis after 2 years) at the individual subject level. Prediction models were trained and cross-validated in a sample of 643 patients with current MDD (2-year chronicity n = 318) and subsequently tested for performance in 161 MDD individuals (2-year chronicity n = 79). Proteomics data showed best unimodal data predictions (AUROC = 0.68). Adding proteomic to clinical data at baseline significantly improved 2-year MDD chronicity predictions (AUROC = 0.63 vs AUROC = 0.78, p = 0.013), while the addition of other -omics data to clinical data did not yield significantly increased model performance. SHAP and enrichment analysis revealed proteomic analytes involved in inflammatory response and lipid metabolism, with fibrinogen levels showing the highest variable importance, followed by symptom severity. Machine learning models outperformed psychiatrists ability to predict two-year chronicity (balanced accuracy = 71% vs 55%). This study showed the added predictive value of combining proteomic, but not other -omic data, with clinical data. Adding other -omic data to proteomics did not further improve predictions. Our results reveal a novel multimodal signature of MDD chronicity that shows clinical potential for individual MDD disease course predictions from baseline measurements.

4
Prognostic predictions in psychosis: exploring the complementary role of machine learning models

van Dee, V.; Kia, S. M. M.; Fregosi, C.; Swildens, W. E.; Alkema, A.; Batalla, A.; van den Berg, C.; Coric, D.; van Dellen, E.; Dijkstra, L. G.; van den Doel, A.; Dominicus, L. S.; Enterman, J.; Gerritse, F.; van der Horst, M. Z.; van Houwelingen, F.; Koch, C. S.; Koomen, L. E. M.; Kromkamp, M.; Lancee, M.; Mouthaan, B. E.; van Rappard, D. F.; Regeer, E. J.; Salet, R. W. J.; Somers, M.; Straalman, J.; de Vette, M. H. T.; Voogt, J.; Winter - van Rossum, I.; Kahn, R. S.; Cahn, W.; Schnack, H. G.

2025-02-02 psychiatry and clinical psychology 10.1101/2025.01.30.25321382 medRxiv
Top 0.1%
7.8%
Show abstract

BACKGROUNDPredicting outcomes in schizophrenia spectrum disorders is challenging due to the variability of individual trajectories. While machine learning (ML) shows promise in outcome prediction, is has not yet been integrated into clinical practice. Understanding how ML models (MLMs) can complement psychiatrists predictions and bridge the gap between MLM capabilities and practical use is key. OBJECTIVEThis study aims to compare the performance of psychiatrists and MLMs in predicting short-term symptomatic and functional remission in patients with first-episode psychosis and explore whether MLMs can improve psychiatrists prognostic accuracy. METHODTwenty-four psychiatrists predicted symptomatic and functional remission probabilities based on written baseline information from 66 patients in the OPTiMiSE trial. ML-generated predictions were then shared with psychiatrists, allowing them to adjust their estimates. A questionnaire assessed trust in MLMs, perceived information gaps, and psychiatrists self-assessed predictive accuracy, which was compared to actual accuracy. FINDINGSThe predictive accuracy of the MLM was comparable to that of psychiatrists for symptomatic remission (MLM: 0.50, psychiatrists: 0.52) and functional remission (MLM: 0.72, psychiatrists: 0.79). Interrater agreement was low but comparable for psychiatrists and the MLM. Although the MLM did not improve overall predictive accuracy, it showed potential in aiding psychiatrists with difficult-to-predict cases. However, psychiatrists struggled to recognize when to rely on the models output and we were unable to determine a clear pattern in these cases based on their characteristics. Psychiatrists could not reliably estimate their predictive accuracy. Psychiatrists expressed moderate to high trust in MLMs for prognostic prediction, but highlighted concerns about the lack of transparency and interpretability of model outputs. CONCLUSIONSMLMs are a promising tool for supporting psychiatric decision-making, particularly in challenging cases. However, their potential remains underutilized due to limitations in predictive accuracy and a lack of clarity in how predictions are generated. Addressing these issues is essential to build trust and foster integration into clinical practice. CLINICAL IMPLICATIONSMLMs are best suited as supplementary tools, providing a second opinion while psychiatrists retain decision-making autonomy. Integrating predictions from both sources may help reduce individual biases and improve accuracy. This approach leverages the strengths of MLMs without compromising clinical responsibility. SUMMARY BOXO_ST_ABSWhat is already known on this topicC_ST_ABSWhile machine learning models (MLMs) show promise in predicting outcomes in psychotic disorders, they have yet to be integrated into clinical practice. Evidence on the predictive accuracy of psychiatrists for these disorders is limited, with only two small studies published before 1990 suggesting moderate accuracy. Comparisons of MLMs and psychiatrists in this context have not been previously conducted. What this study addsThis is the first study to compare the predictive accuracy of psychiatrists with that of an MLM for psychotic disorders and to assess whether an MLM can enhance psychiatrists performance. It highlights that while MLMs do not improve overall accuracy, they may support psychiatrists in difficult cases. Insights into psychiatrists trust in MLMs and the challenges of implementing these models are also provided. How this study might affect research, practice, or policyThe findings emphasize the need for advancements in MLM accuracy, interpretability, and strategies to identify cases where MLMs are most beneficial. These improvements could foster effective integration of MLMs as supplementary tools in clinical practice, aiding psychiatrists in decision-making while maintaining their autonomy.

5
Diagnostic Accuracy and Clinical Reasoning of Multiple Large Language Models in Psychiatry

Jin, K. W.; Rostam-Abadi, Y.; Chaudhary, P.; Garrett, M. A.; Huang, A. S.; Montelongo, M.; Nagpal, C.; Shei, J.; Weathers, J.; Zhang, J. S.; Chen, Q.; Kim, J.; Malgaroli, M.; Mathis, W. S.; Rodriguez, C. I.; Selek, S.; Sharma, M. S.; Pittenger, C.; Yip, S. W.; Zaboski, B. A.; Xu, H.

2026-02-09 psychiatry and clinical psychology 10.64898/2026.02.03.26345402 medRxiv
Top 0.1%
7.7%
Show abstract

ImportanceLarge language models (LLMs) have demonstrated diagnostic potential in several medical specialties, but their application to psychiatry - where diagnosis relies heavily on clinical judgment, narrative interpretation, and reasoning under uncertainty - remains insufficiently evaluated. ObjectiveTo evaluate diagnostic accuracy and clinician-judged reasoning quality of multiple large language models using psychiatric case vignettes. DesignMixed-methods evaluation study of diagnostic accuracy across four LLMs using 196 psychiatric case vignettes (135 published and 61 novel). Clinical reasoning quality was evaluated on a randomly selected subset of 30 vignettes using structured clinician ratings along two reasoning dimensions. The highest-performing model was illustratively compared with psychiatry trainees on the same subset. Diagnostic correctness for the full vignette set was assessed by a separate adjudicator LLM. SettingPublicly available model interfaces, December 2025. ParticipantsFive board-certified psychiatrists evaluated model-generated clinical reasoning. Two psychiatry residents served as the illustrative human comparison. Main Outcomes and MeasuresDiagnostic accuracy and clinician-rated clinical reasoning quality. Diagnostic accuracy was assessed using top-1 accuracy, top-5 accuracy, recall@5, and mean reciprocal rank based on ranked lists of five differential diagnoses per vignette. Clinical reasoning quality was assessed using two 5-point Likert scales adapted from the American Council of Graduate Medical Education Psychiatry Residency Milestones, evaluating data extraction and diagnostic reasoning. ResultsAcross 196 psychiatric case vignettes, Claude Opus 4.5 (Anthropic) achieved the highest diagnostic accuracy (top-1 accuracy, 0.638; top-5 accuracy, 0.801; recall@5, 0.731; mean reciprocal rank, 0.710) and clinician-rated reasoning scores. Higher clinician-rated diagnostic reasoning quality was strongly associated with diagnostic correctness in mixed-effects logistic regression analyses ({beta} = 1.80; p < 0.001), corresponding to an approximately six-fold increase in odds of a correct diagnosis per 1-point increase in reasoning score. In an illustrative comparison, diagnostic accuracy of Claude Opus 4.5 fell within the range observed for psychiatry trainees. Conclusions and RelevanceLLMs demonstrated high diagnostic accuracy and generated clinical reasoning that clinicians judged to be largely coherent and safe. Diagnostic reasoning quality was more strongly associated with diagnostic correctness than data extraction quality, underscoring the importance of evaluating reasoning alongside accuracy when assessing LLMs for clinical decision support in psychiatry. Key PointsO_ST_ABSQuestionC_ST_ABSCan multiple large language models accurately diagnose psychiatric conditions and generate diagnostic reasoning that clinicians judge as coherent, safe, and clinically meaningful? FindingsAcross 196 psychiatric case vignettes, four large language models demonstrated high diagnostic accuracy. In a clinician-evaluated subset of 30 vignettes, model diagnostic accuracy fell within the range observed for psychiatry residents. Clinicians judged model-generated diagnostic reasoning to be largely coherent and safe. Higher clinician-rated reasoning quality was strongly associated with diagnostic correctness, independent of data extraction quality. MeaningEvaluating diagnostic reasoning, in addition to accuracy, may be important when assessing large language models for potential clinical decision support in psychiatry.

6
Detection of Suicidality Through Privacy-Preserving Large Language Models

Wiest, I. C.; Verhees, F. G.; Ferber, D.; Zhu, J.; Bauer, M.; Lewitzka, U.; Pfennig, A.; Mikolas, P.; Kather, J. N.

2024-03-08 psychiatry and clinical psychology 10.1101/2024.03.06.24303763 medRxiv
Top 0.1%
7.7%
Show abstract

ImportanceAttempts to use Artificial Intelligence (AI) in psychiatric disorders show moderate success, high-lighting the potential of incorporating information from clinical assessments to improve the models. The study focuses on using Large Language Models (LLMs) to manage unstructured medical text, particularly for suicide risk detection in psychiatric care. ObjectiveThe study aims to extract information about suicidality status from the admission notes of electronic health records (EHR) using privacy-sensitive, locally hosted LLMs, specifically evaluating the efficacy of Llama-2 models. Main Outcomes and MeasuresThe study compares the performance of several variants of the open source LLM Llama-2 in extracting suicidality status from psychiatric reports against a ground truth defined by human experts, assessing accuracy, sensitivity, specificity, and F1 score across different prompting strategies. ResultsA German fine-tuned Llama-2 model showed the highest accuracy (87.5%), sensitivity (83%) and specificity (91.8%) in identifying suicidality, with significant improvements in sensitivity and specificity across various prompt designs. Conclusions and RelevanceThe study demonstrates the capability of LLMs, particularly Llama-2, in accurately extracting the information on suicidality from psychiatric records while preserving data-privacy. This suggests their application in surveillance systems for psychiatric emergencies and improving the clinical management of suicidality by improving systematic quality control and research. Key PointsO_ST_ABSQuestionC_ST_ABSCan large language models (LLMs) accurately extract information on suicidality from electronic health records (EHR)? FindingsIn this analysis of 100 psychiatric admission notes using Llama-2 models, the German fine-tuned model (Emgerman) demonstrated the highest accuracy (87.5%), sensitivity (83%) and specificity (91.8%) in identifying suicidality, indicating the models effectiveness in on-site processing of clinical documentation for suicide risk detection. MeaningThe study highlights the effectiveness of LLMs, particularly Llama-2, in accurately extracting the information on suicidality from psychiatric records, while preserving data privacy. It recommends further evaluating these models to integrate them into clinical management systems to improve detection of psychiatric emergencies and enhance systematic quality control and research in mental health care.

7
Development and Validation of Machine Learning-Based Prediction of Depression Progression Using EHR Data: A Multi-Institutional Retrospective Cohort Study

Ahadian, P.; Fragano, A.; Guan, T.; Guan, Q.; Shalhout, S. Z.

2025-12-01 psychiatry and clinical psychology 10.1101/2025.11.28.25341207 medRxiv
Top 0.1%
7.6%
Show abstract

BackgroundDepression is a leading cause of global disability. Timely identification of patients at risk for clinical worsening remains a major challenge. Electronic health records (EHRs) facilitate large-scale, real-world analyses of disease trajectories. However, standardized symptom scale data such as the Patient Health Questionnaire-9 are often unavailable or recorded only as unstructured text. In this context, International Classification of Diseases (ICD10) diagnostic-code based severity progression provides a pragmatic alternative for developing predictive tools to identify worsening depression. ObjectiveWe aim to develop and evaluate machine-learning and deep-learning models for predicting ICD10-defined progression from mild to moderate/severe depression using EHR data curated by the MedStar Health Research Institute (MHRI). MethodsWe conducted a multi-institutional retrospective cohort analysis using the MHRI EHR database, which integrates data from 10 hospitals and 300 outpatient sites across the mid-Atlantic. Adults ([&ge;]18 years) with an initial ICD10 diagnosis of mild depression between 2017 and 2023 were included (N=2131). Nonprogressors were defined as patients whose mild major depressive disorder remained mild for 24 months (N=270). Progressors were defined as patients who developed moderate or severe ICD10 depression within 24 months of the index diagnosis (N=533). Data were stratified and split into (60%) training, (20%) validation, and (20%) test subsets. A heterogeneous feature set spanning demographics, healthcare utilization, socioeconomic indices, diagnostic context, and laboratory measurements were available. Logistic regression utilized elastic net regularization with fivefold cross validation, and random forest hyperparameters were tuned by grid search. XGBoost, CatBoost, and a deep neural network (DNN) were trained with standard learning rate, depth, class weighting, and early stopping. A deterministic top model selection framework applied prespecified thresholds of sensitivity at least 0.70 and AUC at least 0.70, and composite rankings integrated accuracy, sensitivity, specificity, and the overfitting gap. ResultsThe analytic cohort included 803 patients with complete two-year follow-up. Under the selection criteria, the DNN failed to meet the AUC threshold (0.671) and was excluded. Among the remaining models, XGBoost achieved the top composite score (accuracy = 0.72; AUC = 0.776; sensitivity = 0.77; specificity = 0.63; overfit gap = 0.112). Logistic regression ranked second (accuracy = 0.71; AUC = 0.797; sensitivity = 0.79; specificity = 0.61; overfit gap = 0.052), followed by CatBoost and random forest, the latter penalized for overfitting (gap = 0.278). The TinyLlama audit note, generated through a local Hugging Face pipeline, confirmed XGBoost as the most balanced model. ConclusionsUsing EHR data from a multi-institutional regional health system, we developed and validated machine-learning models that predicted progression of depression. XGBoost demonstrated the most reliable composite performance. These findings support the feasibility of leveraging socioeconomic and EHR data to predict worsening depression and emphasize the importance of transparent model-selection frameworks for trustworthy clinical artificial intelligence.

8
clickBrick Prompt Engineering: Optimizing Large Language Model Performance in Clinical Psychiatry

Verhees, F. G.; Huth, F.; Meyer, V.; Wolf, F.; Bauer, M.; Pfennig, A.; Ritter, P.; Kather, J. N.; Wiest, I. C.; Mikolas, P.

2025-06-30 psychiatry and clinical psychology 10.1101/2025.06.28.25330267 medRxiv
Top 0.1%
7.2%
Show abstract

BackgroundPrompt engineering has the potential to enhance large language models (LLM) ability to solve tasks through improved in-context learning. In clinical research, the use of LLMs has shown expert-level performance for a variety of tasks ranging from pathology slide classification to identifying suicidality. We introduce clickBrick, a modular prompt-engineering framework, and rigorously test its effectiveness. MethodsHere, we explore the effects of increasingly structuring prompts with the clickBrick framework for a comprehensive psychopathological assessment of 100 index patients from psychiatric electronic health records. We compare the performance of a locally-run LLM (Llama-3.1-70B-Instruct) against an expert-labelled ground truth for a variety of successively built-up prompts for the extraction of 12 transdiagnostic psychopathological criteria. Potential clinical value was explored by training linear support vector machines on outputs from the strongest and weakest prompts to predict discharge ICD-10 main diagnoses for a historical sample of 1,692 patients. OutcomesWe could reliably extract information across 12 distinct psychopathological classification tasks from unstructured clinical text with balanced accuracies spanning 71 % to 94%. Across tasks, we observed a substantially improved extraction accuracy (between +19% and +36%) using clickBrick. The comparison unveiled great variations between prompts with a reasoning prompt performing best in 7 out of 12 domains. Clinical value and internal validity were approximated by downstream classification of eventual psychiatric diagnoses for 1,692 patients. Here, clickBrick led to an improvement in overall classification accuracy from 71% to 76%. InterpretationClickBrick prompt engineering, i.e. iterative, expert-led design and testing, is critical for unlocking LLMs clinical potential. The framework offers a reproducible pathway for deploying trustworthy generative AI across mental health and other clinical fields. FundingThe German Ministry of Research, Technology and Space and the German Research Foundation. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed/MEDLINE articles without language restrictions published before June 25 2025 that combined three concept blocks - "prompt engineering" or related synonyms, "large language model/LLM" or specific model names (e.g., ChatGPT, GPT-4, LLaMA), and psychiatric or mental-health terms (e.g., psychiatry, psychotherapy, depression, anxiety). Additionally, we asked ChatGPT o3 to design and execute a systematic review strategy to also capture not-yet peer-reviewed but relevant pre-prints, given only our manuscript title. After manual de-duplication and abstract screening, three out of 23 identified studies did offer at least some information on their prompting strategies and were conducted on real-world clinical data from psychotherapy transcripts (one study on multi-dimensional counselling therapy, no peer review), or on online patient portal queries (two peer-reviewed studies on (a) empathy evaluation, and (b) provider satisfaction and use of generated response with partial integration with electronic health records. Neither systematically structured their prompts in a transparent way, nor tested reasoning prompts. Beyond psychiatry, one study analyzing automated echocardiography reports did employ a comparison between two different prompts and an expert-led design strategy. A single study used structured and transparent prompt engineering to generate automated responses for simulated problem-solving therapy sessions. None of the highlighted studies reported both head-to-head comparisons of competing prompt strategies for full reproducibility, and their application in real-world care, e.g. on electronic health records. Collectively, the existing literature suggests growing interest but reveals a paucity of rigorous evidence on how prompt engineering impacts large language model performance in clinical psychiatry, particularly in real-world settings. Added value of this studyWe demonstrate reliable information extraction from electronic health records across 12 distinct psychopathological classification tasks from unstructured clinical text and substantially improved extraction accuracy (between +19% and +36%) using clickBrick, our prompt engineering framework. The rationale for such an approach is justified by the surprising identification of zero-shot, few-shot and reasoning prompts as the best performing prompts for different tasks, while a Chain-of-Thought reasoning prompt performs best in 7 out of 12 tasks. And while most studies rely on proprietary language models like openAIs ChatGPT, our locally run version of a popular open-weight model (Llama-3.1-70B-Instruct) allows for privacy safeguarding of sensitive patient data, which is essential for ethical clinical application. Implications of all the available evidenceGenerative artificial intelligence is poised to benefit psychiatric patients greatly, powering advances from therapy delivery to decision support and patient outreach. Rigorous prompt engineering with tools like clickBrick heightens their reliability and credibility, making clickBrick a cornerstone for bringing AI into everyday psychiatric care.

9
Predicting clozapine initiation among patients with schizophrenia via machine learning trained on electronic health record data

Perfalk, E.; Damgaard, J. G.; Danielsen, A. A.; Ostergaard, S. D.

2026-04-20 psychiatry and clinical psychology 10.64898/2026.04.17.26351083 medRxiv
Top 0.1%
7.1%
Show abstract

Background and HypothesisClozapine is the only medication with proven efficacy for treatment-resistant schizophrenia, yet many patients experience delays of several years before initiation. Our aim was to develop and validate a dynamic prediction model for clozapine initiation among patients with schizophrenia trained solely on electronic health record (EHR) data from routine clinical practice. Study DesignEHR data from all adults ([&ge;] 18 years) with a schizophrenia (ICD10: F20) or schizoaffective disorder (ICD10: F25) diagnosis who had been in contact with the Psychiatric Services of the Central Denmark Region between 1 January 2013 and 1 June 2024 were retrieved. 179 structured predictors were engineered (covering, e.g.,diagnoses, medications, coercive measures) and 750 predictors derived from clinical notes. At every psychiatric hospital visit, we predicted if an incident clozapine prescription occured within the next 365 days. XGBoost and logistic regression models were trained on 85% of the data with 5-fold stratified cross-validation. Performance was evaluated on the remaining 15% of the data (held out) using the area under the receiver operating characteristic curve (AUROC). Study ResultsThe training/test set comprised of 194,234/35,527 hospital visits, distributed on 4928/878 unique patients. In the test set, the best XGBoost model achieved an AUROC of 0.81, sensitivity of 32%, positive predictive value of 23% at a 7.5% predicted positive rate. ConclusionsA dynamic prediction model based solely on EHR data predicts clozapine initiation with high discrimination. If implemented as a clinical decision support tool, this model may guide clinicians towards more timely initiation of clozapine treatment.

10
Identifying Psychosis Episodes in Psychiatric Admission Notes via Rule-based Methods, Machine Learning, and Pre-Trained Language Models

Hua, Y.; Blackley, S. V.; Shinn, A. K.; Skinner, J. P.; Moran, L. V.; Zhou, L.

2024-03-19 psychiatry and clinical psychology 10.1101/2024.03.18.24304475 medRxiv
Top 0.1%
6.7%
Show abstract

Early and accurate diagnosis is crucial for effective treatment and improved outcomes, yet identifying psychotic episodes presents significant challenges due to its complex nature and the varied presentation of symptoms among individuals. One of the primary difficulties lies in the underreporting and underdiagnosis of psychosis, compounded by the stigma surrounding mental health and the individuals often diminished insight into their condition. Existing efforts leveraging Electronic Health Records (EHRs) to retrospectively identify psychosis typically rely on structured data, such as medical codes and patient demographics, which frequently lack essential information. Addressing these challenges, our study leverages Natural Language Processing (NLP) algorithms to analyze psychiatric admission notes for the diagnosis of psychosis, providing a detailed evaluation of rule-based algorithms, machine learning models, and pre-trained language models. Additionally, the study investigates the effectiveness of employing keywords to streamline extensive note data before training and evaluating the models. Analyzing 4,617 initial psychiatric admission notes (1,196 cases of psychosis versus 3,433 controls) from 2005 to 2019, we discovered that the XGBoost classifier employing Term Frequency-Inverse Document Frequency (TF-IDF) features derived from notes pre-selected by expert-curated keywords, attained the highest performance with an F1 score of 0.8881 (AUROC [95% CI]: 0.9725 [0.9717, 0.9733]). BlueBERT demonstrated comparable efficacy an F1 score of 0.8841 (AUROC [95% CI]: 0.97 [0.9580, 0.9820]) on the same set of notes. Both models markedly outperformed traditional International Classification of Diseases (ICD) code-based detection methods from discharge summaries, which had an F1 score of 0.7608, thus improving the margin by 0.12. Furthermore, our findings indicate that keyword pre-selection markedly enhances the performance of both machine learning and pre-trained language models. This study illustrates the potential of NLP techniques to improve psychosis detection within admission notes and aims to serve as a foundational reference for future research on applying NLP for psychosis identification in EHR notes.

11
AI-Powered Triage of Suicidal Ideation in Adolescents: A Comparative Evaluation of Large Language Models Using Synthetic Clinical Vignettes

Mansoor, M. A.

2025-08-07 psychiatry and clinical psychology 10.1101/2025.08.05.25333046 medRxiv
Top 0.1%
6.6%
Show abstract

ObjectiveTo evaluate the performance of leading Large Language Models (LLMs) in classifying suicide risk and generating clinically appropriate action plans for adolescent psychiatric cases presented through synthetic clinical vignettes. MethodsWe developed 40 synthetic clinical vignettes depicting adolescents with varying levels of suicide risk, structured according to established clinical formulation principles. A gold standard for risk level, based on the Columbia-Suicide Severity Rating Scale (C-SSRS) framework, and corresponding clinical actions was established for each vignette by a panel of two board-certified child and adolescent psychiatrists. Three LLMs (GPT-4o, Claude 3.5 Sonnet, Llama-3.1-70B) were prompted using a structured chain-of-thought methodology to classify risk and propose a detailed action plan. Performance was assessed using quantitative classification metrics (accuracy, precision, recall, F1-score) and qualitative thematic analysis of the generated action plans. ResultsQuantitative analysis of risk classification revealed variable performance. GPT-4o achieved the highest accuracy (82.5%), followed by Claude 3.5 Sonnet (75.0%) and Llama-3.1-70B (67.5%). F1-scores demonstrated challenges in correctly identifying higher-risk categories, particularly for nuanced presentations of intent. Qualitative thematic analysis of the action plans identified consistent adherence to basic safety protocols (e.g., recommending emergency evaluation for high-risk cases). However, significant and critical failures were pervasive, including the omission of crucial inquiries about access to lethal means, failure to incorporate protective factors into planning, and the generation of clinically inappropriate therapeutic reassurance in a triage context. ConclusionsWhile LLMs demonstrate a nascent ability to process clinical information for suicide risk assessment, significant deficits in clinical reasoning and safety planning persist. Their performance on idealized synthetic data suggests these models are not yet suitable for autonomous clinical decision-making. These findings underscore the imperative for rigorous, clinically-grounded evaluation frameworks and the development of human-in-the-loop systems to ensure patient safety in any future deployment. Key MessagesO_ST_ABSWhat is already known on this topicC_ST_ABSSuicide is a leading cause of death in adolescents, yet current clinical risk assessment tools and subjective judgments have limited predictive accuracy and are difficult to scale in the face of rising demand and workforce shortages. What this study addsThis study provides a direct comparative evaluation of multiple state-of-the-art Large Language Models on a standardized adolescent suicide risk triage task, using a synthetic data methodology that allows for controlled assessment of both classification accuracy and the clinical appropriateness of generated action plans. How this study might affect research, practice or policyThe findings highlight the potential of LLMs as adjunctive tools in non-specialist settings but also reveal critical safety and reliability gaps that must be addressed through further research, the development of ethical guidelines, and regulatory oversight before any clinical implementation can be considered.

12
Cross-site predictions of readmission after psychiatric hospitalization with mood or psychotic disorders

Ren, B.; Yoon, W.; Thomas, S.; Savova, G.; Miller, T.; Hall, M.-H.

2024-08-26 psychiatry and clinical psychology 10.1101/2024.08.26.24312586 medRxiv
Top 0.1%
6.5%
Show abstract

Patients with mood or psychotic disorders have high rates of unplanned readmission, and predicting readmission likelihood may guide discharge decisions. In this retrospective, multi-site study, we assess the predictive power of various structured variables from electronic health records for all-cause readmission in each site separately and evaluate the generalizability of the in-site prediction models across sites. We find that the set of relevant predictors vary significantly across. For example, length of stay is strongly predictive of readmission at only three out of the four sites. We also find a general lack of cross-site generalizability of the in-site prediction models, with in-site predictions having an average F1 score of 0.666, compared to an average F1 score of 0.551 for cross-site predictions. The generalizability cannot be improved even after adjusting for differences in the distributions of predictors. These results indicate that, with this set of predictors, fitting individual models at each site is necessary to achieve reasonable prediction accuracy. Additionally, they suggest that more sophisticated predictors variables or predictive algorithms are needed to develop generalizable models capable of extracting robust insights into the root causes of early psychiatric readmissions.

13
Dynamical instability measured by temporal entropy improves psychiatric classification across cohorts

Shoji, T.; Nakaki, R.

2026-04-30 bioinformatics 10.64898/2026.04.28.721265 medRxiv
Top 0.1%
6.3%
Show abstract

Psychiatric disorders, such as attention-deficit/hyperactivity disorder, autism spectrum disorder, and schizophrenia, are clinically heterogeneous and lack objective biomarkers for reliable diagnosis. Although blood transcriptomic data have been proposed as a potential source of diagnostic information, their generalizability across independent cohorts remains unclear. This study aimed to assess whether biologically informed measures of dynamic instability enhance the reproducibility and generalizability of psychiatric classifications based on peripheral blood data by integrating publicly available blood transcriptomic datasets from multiple cohorts and evaluating classification performance using individual-level cross-validation and study-level holdout validation. To investigate the underlying biological structure, we applied a dynamic systems framework, including pseudotime-based vector field inference and attractor analysis. Additionally, we introduced temporal entropy as a measure of dynamic instability in the inferred transcriptomic trajectories. High classification performance was observed in individual-level cross-validation (area under the receiver operating characteristic [AUROC] > 0.8 across several comparisons); however, performance decreased substantially in study-level validation (AUROC {approx} 0.5-0.7), indicating limited generalizability. Attractor analysis revealed that transcriptomic states formed continuous and overlapping structures rather than distinct diagnostic clusters. Stratification based on temporal entropy identified a subset of individuals with unstable transcriptomic dynamics, and excluding these individuals improved the classification performance across most diagnostic pairs (AUROC > 0.7). These findings suggest that transcriptomic variability and dynamic instability contribute to the limited reproducibility of psychiatric classifications. Incorporating temporal entropy as a measure of system-level instability may enhance the robustness and interpretability of biomarker-based models and provide a new perspective on psychiatric disorders as dynamic systems.

14
A Local Outpatient Practice-Level Prediction Model for Short-Term Psychiatric Emergency Presentation

Havlik, J. L.; Tyrrell, B.; Bell, N.; Polaschek, J.; Arzubi, E. R.

2026-07-01 psychiatry and clinical psychology 10.64898/2026.06.29.26356785 medRxiv
Top 0.1%
6.1%
Show abstract

Importance: Psychiatric emergency department (ED) presentations are difficult to predict using general medical risk stratification tools. Health information exchange (HIE) data may improve prediction by capturing fragmented care across settings. Objective: To develop and temporally validate a machine learning model using HIE and geospatial data to predict 30-day psychiatric ED presentation among outpatients receiving psychiatric care and to compare its performance with standard clinical risk scores. Design, Setting, and Participants: This retrospective cohort study included patients seen at Frontier Psychiatry with records in the Big Sky Care Connect statewide HIE. Structured clinical data were linked to zip code-level sociodemographic measures. The analytic unit was the patient snapshot, defined as all structured data available up to a given point. Models were evaluated in temporally separated train and test sets. Exposures: Predictors derived from HIE structured data, including prior utilization, diagnoses, medications, laboratory data, and zip code-linked geospatial deprivation and vulnerability measures. Main Outcomes and Measures: The primary outcome was psychiatric ED presentation within 30 days, identified from structured encounter-type fields and primary diagnosis codes for psychiatric or substance use disorders. Model discrimination was compared with a parsimonious clinical baseline model and LACE and Elixhauser scores. Results: In the test set, 343 of 16,469 snapshots (2.1%) were followed by a qualifying psychiatric ED presentation within 30 days, corresponding to 102 ED visits among 68 patients. The machine learning model showed discrimination in temporally held-out testing and outperformed the clinical baseline model as well as LACE and Elixhauser scores. At a prespecified decision threshold, the model reduced the number needed to evaluate from more than 40 with universal screening to 3.4 to identify 1 true-positive case, while identifying over two fifths of 30-day psychiatric ED presentations. Conclusions and Relevance: In this retrospective cohort study, a locally developed machine learning model using statewide HIE data showed improved prediction of 30-day psychiatric ED presentation compared with selected general-purpose risk scores. The results support the feasibility of HIE-enabled local psychiatric risk modeling and suggest other practices could develop similarly tailored models. Prospective studies are needed to assess clinical utility and effects on outcomes.

15
Phenotyping Antidepressant Treatment Response with Deep Learning in Electronic Health Records

Sheu, Y.-h.; Magdamo, C.; Miller, M.; Das, S.; Blacker, D.; Smoller, J. W.

2021-08-04 psychiatry and clinical psychology 10.1101/2021.08.04.21261512 medRxiv
Top 0.1%
6.1%
Show abstract

Efficient, accurate phenotyping for antidepressant treatment response in electronic health records (EHRs) could facilitate precision psychiatry applications but remains a challenge. Increasingly, artificial intelligence methods using "deep learning" applied to clinical data have shown promise in complex classification problems. Here, we systematically evaluate the performance of eight deep-learning-based natural language processing models in classifying response to antidepressants in a large real-world healthcare setting. We obtained data spanning 1990-2018 for adults with depression and a co-occurring antidepressant prescription from the EHR data warehouse of the Mass General Brigham healthcare system (n=111,572). Clinical notes were collected for the following time windows after antidepressant initiation: (1) 2 days to 4 weeks, (2) 4-12 weeks, and (3) 12-26 weeks. A stratified random sample of these note sets (total 4,299 across time periods) were manually reviewed to classify response status as "improved" or "no evidence of improvement" in depression symptoms. All models performed well, with areas under the receiver operator curve (AUROC) of at least 0.80. Positive predictive values (PPVs) ranged from 0.72 - 0.91. In general, models incorporating more information-dense and longer text sequences performed better than others. The best performing model (Longformer-large with sliding window) had an AUROC = 0.88 and PPV = 0.84 at a specificity of 0.88. Our results indicate that deep learning methods applied to EHR data can accurately classify antidepressant response in a real-world healthcare setting. Automated treatment response classification may facilitate a range of research and clinical decision support applications.

16
Machine Learning Models for the Prediction of Early-Onset Bipolar Using Electronic Health Records

Wang, B.; Sheu, Y.-h.; Lee, H.; Mealer, R. G.; Castro, V. M.; Smoller, J. W.

2024-02-21 psychiatry and clinical psychology 10.1101/2024.02.19.24302919 medRxiv
Top 0.1%
6.0%
Show abstract

ObjectiveEarly identification of bipolar disorder (BD) provides an important opportunity for timely intervention. In this study, we aimed to develop machine learning models using large-scale electronic health record (EHR) data including clinical notes for predicting early-onset BD. MethodStructured and unstructured data were extracted from the longitudinal EHR of the Mass General Brigham health system. We defined three cohorts aged 10 - 25 years: (1) the full youth cohort (N=300,398); (2) a sub-cohort defined by having a mental health visit (N=105,461); (3) a sub-cohort defined by having a diagnosis of mood disorder or ADHD (N=35,213). By adopting a prospective landmark modeling approach that aligns with clinical practice, we developed and validated a range of machine learning models including neural network-based models, across different cohorts and prediction windows. ResultsWe found the two tree-based models, Random forests (RF) and light gradient-boosting machine (LGBM), achieving good discriminative performance across different clinical settings (area under the receiver operating characteristic curve 0.76-0.88 for RF and 0.74-0.89 for LGBM). In addition, we showed comparable performance can be achieved with a greatly reduced set of features, demonstrating computational efficiency can be attained without significant compromise of model accuracy. ConclusionGood discriminative performance for early-onset BD is achieved utilizing large-scale EHR data. Our study offers a scalable and accurate method for identifying youth at risk for BD that could help inform clinical decision making and facilitate early intervention. Future work includes evaluating the portability of our approach to other healthcare systems and exploring considerations regarding possible implementation.

17
A longitudinal study of depressive symptom trajectories and risk factors in congestive heart failure

Gallucci, J.; Ng, J.; Secara, M. T.; Jones, B. D. M.; Hawco, C.; Husain, M. O.; Husain, N.; Chaudhry, I. B.; Voineskos, A. N.; Husain, M. I.

2024-09-17 cardiovascular medicine 10.1101/2024.09.16.24313783 medRxiv
Top 0.1%
5.6%
Show abstract

BackgroundDepression is prevalent among patients with congestive heart failure (CHF) and is associated with increased mortality and healthcare utilization. However, most research has focused on high-income countries, leaving a gap in knowledge regarding the relationship between depression and CHF in low-to-middle-income countries (LMICs). This study aimed to delineate depressive symptom trajectories and identify potential risk factors for poor outcomes among CHF patients. MethodsLongitudinal data from 783 patients with CHF from public hospitals in Karachi, Pakistan was analyzed. Depressive symptom severity was assessed using the Beck Depression Inventory (BDI). Baseline and 6-month follow-up BDI scores were clustered through Gaussian Mixture Modeling to identify distinct depressive symptom subgroups and extract trajectory labels. Further, a random forest algorithm was utilized to determine baseline demographic, clinical, and behavioral predictors for each trajectory. ResultsFour depressive symptom trajectories were identified: good prognosis, remitting course, clinical worsening, and persistent course. Risk factors associated with persistent depressive symptoms included lower quality of life and the New York Heart Association (NYHA) class 3 classification of CHF. Protective factors linked to a good prognosis included less disability and a non-NYHA class 3 classification of CHF. ConclusionsBy identifying key characteristics of patients at heightened risk of depression, clinicians can be aware of risk factors and better identify patients who may need greater monitoring and appropriate follow-up care. Clinical Perspective What is new?O_LITo the best of our knowledge, this is the first study to use machine learning techniques to investigate depressive symptom trajectories in CHF patients from an LMIC. C_LIO_LIFour distinct depressive symptom trajectories were identified, ranging from good prognosis to persistent depressive symptoms. C_LIO_LIThis study highlights protective and risk factors associated with these trajectories based on patients demographics and clinical presentations at baseline. C_LI What are the clinical implications?O_LIPersonalized interventions based on identified protective factors for high-risk CHF patients could enhance both mental health and cardiovascular outcomes. C_LIO_LIEarly detection and management of depression, particularly in patients with poor quality of life or advanced heart failure, may help reduce healthcare utilization and mortality. C_LIO_LIThis study emphasizes the importance of routine depression screening in CHF patients, especially in LMICs, to enhance overall patient care and outcomes. C_LI

18
Predicting PANSS symptoms in schizophrenia spectrum disorders using speech only: an international, multi-centre, retrospective, computational study across multiple languages

He, R.; Kirdun, M.; Palominos, C.; Navarrete Orejudo, L.; Barthelemy, S.; Bhola, S.; Ciampelli, S.; Decker, A.; Demirlek, C.; Fusaroli, R.; Garcia-Molina, J. T.; Gimenez, G.; Huppi, R.; Koelkebeck, K.; Lecomte, A.; Qiu, R.; Simonsen, A.; Tourneur, V.; Verim, B.; Wang, H.; Yalincetin, B.; Yin, S.; Zhou, Y.; Amblard, M.; Ayesa Arriola, R.; Bora, E.; de Boer, J.; Figueroa-Barra, A. I.; Koops, S.; Musiol, M.; Palaniyappan, L.; Parola, A.; Spaniel, F.; Tang, S. X.; Sommer, I. E.; Homan, P.; Hinzen, W.

2026-02-28 psychiatry and clinical psychology 10.64898/2026.02.20.26345632 medRxiv
Top 0.1%
5.5%
Show abstract

Backgroundspeech carries cues to variation in mental state in schizophrenia spectrum disorders/psychotic disorders, typically indexed with clinician-rated scales such as the PANSS. Progress in the automation of speech-based symptom modelling has been constrained by data scale and the underrepresentation of low-resource languages. In this study, we aggregate multi-center recordings to assemble a large corpus and assess symptom-prediction models at scale, to enable more objective and efficient assessments and the early detection of relapse-related signals from speech. MethodsWe compiled data from 453 patients with schizophrenia spectrum disorders, recruited from ten global sites, and clipped their speech recordings into 6,664 segments. Across three feature sets, acoustic-prosodic profile, pretrained multilingual embeddings, and their concatenation, we compared 16 algorithms to predict eight relapse-related PANSS items, including three positive (P1, P2, P3), three negative (N1, N4, N6), and two general (G5, G9) items, on speaker-disjoint splits (80% train, 10% test, and 10% validation). Performance was assessed by root-mean-squared-error (RMSE) at both segment and participant (median aggregation) levels. Best model per item underwent bias checks for age, sex, education, and symptom severity. OutcomesBest-performing models predicted symptoms with prediction errors of 1{middle dot}5 PANSS points or lower: P1 1{middle dot}494/1{middle dot}527, P2 1{middle dot}318/1{middle dot}107, P3 1{middle dot}407/1{middle dot}542, N1 1{middle dot}029/1{middle dot}030, N4 1{middle dot}452/1{middle dot}430, N6 0{middle dot}860/0{middle dot}855, G5 0{middle dot}850/0{middle dot}882, G9 1{middle dot}213/1{middle dot}282 (segment/participant). Performance of the pretrained multilingual embeddings surpassed acoustic-prosodic features and their concatenation. Results were comparable in low-resource languages (e.g., Czech). We found no bias by age, sex, or education, aside from reduced N4 accuracy in males; but performance degraded with higher symptom severity. InterpretationSpeech can support automatic assessment of schizophrenia symptoms using pretrained multilingual embeddings, even without the use of transcripts. Such models show promise as clinically meaningful, efficient, and low-burden tools for real-time monitoring of symptom trajectories. FundingEU Horizon research and innovation programme. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSAutomatic assessment of disease severity is a key issue in schizophrenia research, for which spontaneous speech offers a cost-effective, automatable solution. To evaluate existing evidence for speech-based symptom assessment, two reviewers (RHe, MK) searched PubMed, IEEE Xplore, arXiv, bioRxiv, and medRxiv for publications from inception to Aug 25, 2025, using the terms: ("symptom" OR "PANSS" OR "Positive and Negative Syndrome Scale") AND ("psychosis" OR "schizophrenia") AND ("language" OR "speech" OR "spontaneous speech") AND ("prediction" OR "machine learning" OR "deep learning" OR "algorithm" OR "neural network" OR "AI" OR "artificial intelligence"). Fourteen studies on symptom-level modelling were identified. Ten studies dichotomized clinical scores (e.g., PANSS) into low vs high for classification: five used conventional ML (e.g., random forests) and five used neural networks, with F1 scores ranging from 0{middle dot}60-0{middle dot}85. The remaining four studies, and two of the ten studies as mentioned above, modelled raw scores directly as regression tasks. Two relied solely on conventional regressors and the rest used neural networks, with errors from 0{middle dot}487 for single items (scale 1-7) to 8{middle dot}04 for summed scores (scale 18-126). All studies used free speech for elicitation, except one study, which used a reading task. Three studies incorporated additional tasks, such as picture description and immediate recall. None were multilingual: nine were in English, three in Chinese, one in Swiss German, and one in Brazilian Portuguese. Features spanned a wide range, including acoustic-prosodic profiles, morpho-syntactic structure, semantic organization, pragmatics (including sentiments), and even visual features capturing movement during talking. Representations from pretrained language models were also widely employed. Sample sizes (counting patients with schizophrenia) were generally small: eleven studies enrolled <50 patients, one had 65, and only two exceeded 100 patients. Some increased their effective sample size via multiple recordings per patient or by adding healthy controls and/or patients with other psychiatric disorders (e.g., depression). Added value of this studyTo our knowledge, this is the first multilingual, speech-based study modelling schizophrenia symptom severity with machine learning approach, and it includes the largest cohort of patients with schizophrenia to date. We further increased effective sample size by using diverse elicitation tasks and segmenting recordings into clips. This multilingual corpus empowers the usage of complex models and supports transfer learning from high-resource languages (e.g., English) to low-resource ones (e.g., Czech). For each of eight selected relapse-related PANSS items, the best audio-only models achieved RMSE < 1{middle dot}5, underscoring clinical relevance. We assessed potential biases: no effects were found for age, sex, or education (except poorer N4 performance in males), though performance declined at higher symptom severity. Trained models are released for use. Implications of all the available evidenceWe show that speech is a powerful signal for automatic assessment of schizophrenia symptom severity and holds promise for relapse prediction, even without transcripts. The approach readily extends to incorporate textual features (from manual or automatic transcripts) and more advanced models. Prospective studies with repeated recordings across relapse episodes are needed to validate the utility of our models on relapse prediction, for the sake of supporting precision psychiatry while reducing clinician burden.

19
Data Diversity vs. Model Complexity in the Prediction of Pediatric Bipolar Disorder: Evidence from Academic and Community Clinical Samples

Shi, Z.; Youngstrom, E. A.; Liu, Y.; Youngstrom, J. K.; Findling, R. L.

2026-03-27 psychiatry and clinical psychology 10.64898/2026.03.26.26349447 medRxiv
Top 0.1%
5.5%
Show abstract

Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed to enhance diagnostic reliability. We evaluated a clinical decision tool (nomogram), statistical methods (logistic regression, LASSO), machine learning (support vector machine, random forest, k-nearest neighbors, extreme gradient boosting), and deep learning model (multilayer perceptron) for pediatric bipolar disorder prediction across two datasets collected in academic (N=550) and community (N=511) clinical settings. We compared three modeling strategies: cross-dataset validation, cross-dataset with interaction terms, and mixed-dataset. We assessed model performance using discrimination ability, calibration, and predictor importance ranking. In the baseline cross-dataset approach, all models showed good internal discrimination in the academic dataset; but external discrimination in the community dataset substantially declined. Interaction-enhanced models slightly improved internal discrimination but not external performance or calibration. Recalibration prominently improved cross-dataset calibration without compromising discrimination, indicating that transportability problems were largely driven by probability scaling. Models trained on mixed datasets exhibited much stronger external discrimination and calibration. Across models and training strategies, family risk and PGBI-10M were consistently ranked as the most important predictors. Predictive models for pediatric bipolar disorder showed strong internal performance but limited cross-setting generalizability due to dataset shift and miscalibration. Increasing model complexity did not improve external performance, whereas training on pooled data substantially improved both discrimination and calibration. Findings suggest that sampling diversity, rather than model complexity, is more valuable for developing clinically useful and generalizable psychiatric prediction models, underscoring the importance of open and collaborative datasets.

20
Structured large language model extraction of clinical factors from electronic health record text supports scalable psychiatric severity prediction

Stephenson, C.; Camassa, A.; Wagner, M.; Shirazi, A. H.; Alavi, N.; Omrani, M.

2026-05-13 psychiatry and clinical psychology 10.64898/2026.05.11.26352839 medRxiv
Top 0.1%
5.4%
Show abstract

BackgroundMental health systems face escalating demand that exceeds clinician capacity, making accurate severity-based triage a critical bottleneck. Severity assessment guides treatment intensity, resource allocation, and risk management, yet most clinically relevant information remains embedded in unstructured electronic health record (EHR) narratives, limiting its utility for scalable decision support. ObjectivesThis study evaluates whether a single large language model (LLM) can autonomously extract clinical factors from psychiatric EHR narratives, derive predictive weights from those factors, and use the resulting structured representation to predict clinician-implied severity at scale. MethodsFrom a Mayo Clinic repository of more than 2.7 million encounters, 15,000 de-identified psychiatric notes were sampled into a 5,000-patient discovery cohort and a 10,000-patient replication cohort. The same LLM (Llama 3 8B Instruct) extracted 17 background clinical factors and 3 treatment-action factors from each note. Severity reference labels were derived from the treatment-action factors using pre-specified clinical criteria. The LLM independently derived two factor-weight dictionaries from the discovery cohort: one capturing risk-oriented predictors of severe presentations and one capturing protective predictors. Five weighting conditions were then evaluated against the severity labels: the two LLM-derived dictionaries, two controls (LLM-derived variables with randomized weights; clinically irrelevant variables with arbitrary weights), and an unweighted zero-shot baseline. Performance was assessed across 928 valid iterations in the replication cohort. ResultsLLM-derived structured conditions significantly outperformed all controls and the baseline, with statistically equivalent performance between the two structured conditions. Improvements in precision and recall were balanced, indicating gains in discriminative capacity rather than threshold shifts. The variables and weights the LLM derived as predictors of severe presentations aligned closely with established clinical determinants of psychiatric severity. ConclusionA single LLM can derive clinically meaningful factor weights from unstructured EHR narratives and use them to predict psychiatric severity at scale, supporting a viable path toward interpretable, scalable triage in resource-constrained mental health systems.