Development of a natural language processing application to extract and categorize mentions of violence from mental healthcare records text
Li, L.; Sondh, S.; Sondh, H. K.; Stewart, R.; Roberts, A.
Show abstract
BackgroundExperiences of violence are reported frequently by mental health service users, victims of violence are at a greater risk of mental health disorders, and violence may sometimes occur as a consequence of a mental disorder. Electronic health records (EHRs) are an important source of information about healthcare, and its social context. Occurrences of violence are not routinely recorded as structured data in EHRs but are however recorded in the free text narrative. ObjectiveOur objective was to address this research gap by creating a natural language processing (NLP) application that extracts information related to various forms of violence (physical (non-sexual), sexual, emotional, and financial) from the EHR of a large south London mental health service. Additionally, we aimed to extract features concerning the patients role (victimization vs. perpetration), timing (recent vs. historic), domestic context, presence (actual, threat, or unclear), and polarity (affirmed, abstract, or negated) of the violent behaviors. MethodsTwo raters independently annotated 6,500 randomly selected segments of clinical notes containing violence-related keywords from a large mental healthcare provider in South London, each containing 400 characters (with approximately 200 characters before and after the keyword) after rigorous training using a pre-defined and approved coding book provided by senior professionals. We utilized 90% of the annotated data for fine-tuning a multi-label BERT model (employing 5-fold cross-validation) with the remaining 10% of data reserved for a blind test. ResultsThe model performed well on the blind test set for emotional violence (F1= 0.89), financial violence (0.88), physical (non-sexual) violence (0.84), and unspecified violence (0.81), and the patient role (0.89 as perpetrator; 0.84 as victim), polarity (0.89 for affirmed behavior), presence (0.95 for actual violence), and domestic settings (0.88). We were unable to achieve satisfactory results in capturing temporal aspects (0.65 for past violence). ConclusionsWe were able to improve substantially on previously developed NLP for ascertaining violence in routine mental health records, providing novel opportunities for both surveillance and research.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of machine learning algorithms in national data to classify the risk of self-harm among young adults in hospital: a retrospective study 94%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 91%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 93%
- Developing and validating the Nepalese Abuse Assessment Screen (N-AAS) for identifying domestic violence among pregnant women in Nepal 92%
Similar papers in this journal
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 94%
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 92%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 92%
Similar papers in this journal
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 93%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 93%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 91%
Similar papers in this journal
- Comparing human vs. machine-assisted analysis to develop a new approach for Big Qualitative Data Analysis 92%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 92%
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.