Machine Learning in Psychiatric Health Records: A Gold Standard Approach to Trauma Annotation
Atwood, B.; Holderness, E.; Verhagen, M.; Shinn, A. K.; Cawkwell, P.; Cerruti, H.; Pustejovsky, J.; Hall, M.-H.
Show abstract
Psychiatric electronic health records present unique challenges for machine learning due to their unstructured, complex, and variable nature. This study aimed to create a gold standard dataset by identifying a cohort of patients with psychotic disorders and posttraumatic stress disorder, (PTSD), developing clinically-informed guidelines for annotating traumatic events in their health records and to create a gold standard publicly available dataset, and demonstrating the datasets suitability for training machine learning models to detect indicators of symptoms, substance use, and trauma in new records. We compiled a representative corpus of 200 narrative heavy health records (470,489 tokens) from a centralized database and developed a detailed annotation scheme with a team of clinical experts and computational linguistics. Clinicians annotated the corpus for trauma-related events and relevant clinical information with high inter-annotator agreement (0.715 for entity/span tags and 0.874 for attributes). Additionally, machine learning models were developed to demonstrate practical viability of the gold standard corpus for machine learning applications, achieving a micro F1 score of 0.76 and 0.82 for spans and attributes respectively, indicative of their predictive reliability. This study established the first gold-standard dataset for the complex task of labelling traumatic features in psychiatric health records. High inter-annotator agreement and model performance illustrate its utility in advancing the application of machine learning in psychiatric healthcare in order to better understand disease heterogeneity and treatment implications.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 92%
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 94%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 92%
Similar papers in this journal
- Integrating Expert Knowledge into Large Language Models Improves Performance for Psychiatric Reasoning and Diagnosis 95%
- Symptom Monitoring based on Digital Data Collection During Inpatient Treatment of Schizophrenia Spectrum Disorders – a Feasibility Study 91%
- Piloting Forensic Tele-Mental Health Evaluations of Asylum Seekers 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.