Tweets Classification for Digital Epidemiology of Childhood Health Outcomes Using Pre-Trained Language Models
Wickrama Arachchi Athukoralage, D. S.; Atapattu, T.; Thilakaratne, M.; Falkner, K.
Show abstract
This paper presents our approaches for the SMM4H24 Shared Task 5 on the binary classification of English tweets reporting childrens medical disorders. Our first approach involves fine-tuning a single RoBERTa-large model, while the second approach entails ensembling the results of three fine-tuned BERTweet-large models. We demonstrate that although both approaches exhibit identical performance on validation data, the BERTweet-large ensemble excels on test data. Our best-performing system achieves an F1-score of 0.938 on test data, out-performing the benchmark classifier by 1.18%.
Matching journals
The top 14 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 95%
- Use of large language models as a scalable approach to understanding public health discourse 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
Similar papers in this journal
Similar papers in this journal
- Machine Learning Models Predict the Emergence of Depression in Argentinean College Students during Periods of COVID-19 Quarantine 91%
- Deep Multimodal Representations and Classification of First-Episode Psychosis via Live Face Processing 90%
- Optimising a Simple Fully Convolutional Network (SFCN) for accurate brain age prediction in the PAC 2019 challenge 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.