Multimodal Speech and Text Models to Detect Suicidal Risks in Adolescents
Jin, J.; Liu, Y.; Dai, Y.
Show abstract
BackgroundEarly detection of suicide risk in adolescents is crucial but faces challenges including stigma, reluctance to disclose suicidal thoughts, and limited accessibility of mental health resources. Traditional assessment methods may miss at-risk populations, particularly in community settings. This study aimed to explore whether multimodal analysis combining acoustic and linguistic features can improve prediction of suicide risk in adolescents. MethodsVoice recordings and transcribed text from 600 Chinese adolescents (aged 10-18 years) were collected from 47 schools in Guangdong, China. Suicide risk labels were derived from the Mini International Neuropsychiatric Interview for Children and Adolescents (MINI-KID). The dataset included three voice tasks: answering an open-ended question about emotional regulation, reading a standard passage, and describing a face with negative emotions. Features were extracted using pre-trained models (EMOTION2VEC for acoustic features, Paraformer for speech-to-text conversion, and Tongyi Qianwens text-embedding-v3 for text features). We applied various machine learning classifiers including Support Vector Machine, Multi-layer Perceptron, Random Forest, and XGBoost to develop both single-modal and multimodal prediction models. Front-end fusion (FF) and back-end fusion (BF) techniques were employed to combine acoustic and linguistic features. ResultsFusion models combining both acoustic and linguistic features consistently outperformed individual models. The model with both front-end and back-end fusion achieved the highest performance with an accuracy of 0.73, precision of 0.70, recall of 0.80, and F1 score of 0.74. Front-end fusion alone achieved the highest Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.767. Models performed equivalently across age groups but significantly better in females (AUROC = 0.72) compared to males (AUROC = 0.46). ConclusionsMultimodal analysis combining acoustic and linguistic features significantly improves predictive accuracy for adolescent suicide risk detection compared to single-modal approaches. This approach offers a promising method for early identification of at-risk adolescents in community settings, potentially enabling timely intervention. Further external validation with larger samples is needed to optimize these models for clinical application.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 93%
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 92%
- Fear of Infection and Sufficient Vaccine Reservation Information Might Drive Rapid Coronavirus Disease 2019 Vaccination in Japan: Evidence from Twitter Analysis 92%
Similar papers in this journal
- Understanding Psychiatric Illness Through Natural Language Processing (UNDERPIN): Rationale, Design, and Methodology 93%
- Machine Learning Models Predict the Emergence of Depression in Argentinean College Students during Periods of COVID-19 Quarantine 92%
- Passive sensing data predicts stress in university students: A supervised machine learning method for digital phenotyping 91%
Similar papers in this journal
Similar papers in this journal
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 94%
- Off-body Sleep Analysis for Predicting Adverse Behavior in Individuals with Autism Spectrum Disorder 93%
- LncDLSM: Identification of Long Non-coding RNAs with Deep Learning-based Sequence Model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.