Back

Tweets Classification for Digital Epidemiology of Childhood Health Outcomes Using Pre-Trained Language Models

Wickrama Arachchi Athukoralage, D. S.; Atapattu, T.; Thilakaratne, M.; Falkner, K.

2024-06-12 public and global health
10.1101/2024.06.11.24308776 medRxiv
Show abstract

This paper presents our approaches for the SMM4H24 Shared Task 5 on the binary classification of English tweets reporting childrens medical disorders. Our first approach involves fine-tuning a single RoBERTa-large model, while the second approach entails ensembling the results of three fine-tuned BERTweet-large models. We demonstrate that although both approaches exhibit identical performance on validation data, the BERTweet-large ensemble excels on test data. Our best-performing system achieves an F1-score of 0.938 on test data, out-performing the benchmark classifier by 1.18%.

Matching journals

The top 14 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.