Large Language Models for Detecting Body Weight Changes as Side Effects of Antidepressants in User-Generated Online Content
Yokoyama, T.; Natter, J.; Godet, J.
Show abstract
ObjectiveHealthcare websites allow patients to share their experiences with their treatments. Drug testimonials provide useful information for real-world evidence, particularly on the occurrence of side effects that may be underreported. We investigated the potential of large language models (LLMs) for detecting signals of body weight change as under-reported side effect of antidepressants in user-generated online content. Materials and MethodsA database of 8,000 user-generated comments about the 32 FDA-approved antidepressants was collected from healthcare social websites. These comments were manually annotated under the supervision of drug experts. Several pre-trained LLMs derived from BERT were fine-tuned to automatically classify comments describing weight gain, weight loss, or the absence of reference to a weight change. Zero-shot classification was also performed. Performance was evaluated on a test set by measuring the weighted precision, recall, F1-score and the prediction accuracy. ResultsAfter fine-tuning, most of the BERT-derived LLMs showed weighted F1-scores above 97%. LLMs with higher number of parameters used in zero-shot classification almost reached the same performance. The main source of errors in predictions came from situations where the machine predicted falsely weight gain or loss, because the text mentioned these elements but for a different molecule than the one for which the comment was written. ConclusionEven fine-tuned LLMs with limited numbers of parameters showed interesting results for the detection of adverse events from online patient testimonials, suggesting they can be used at scale for real-world evidence.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Which social media platforms facilitate monitoring the opioid crisis? 94%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 92%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Personalized Mood Prediction from Patterns of Behavior Collected with Smartphones 92%
Similar papers in this journal
- Reducing maladaptive behavior in neuropsychiatric disorders using network modification 91%
- Machine Learning Models Predict the Emergence of Depression in Argentinean College Students during Periods of COVID-19 Quarantine 91%
- Computational Psychiatry Research Map (CPSYMAP): a New Database for Visualizing Research Papers 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.