Text Augmentations with R-drop for Classification of Tweets Self-Reporting Covid-19
Francis, S.
Show abstract
This paper presents models created for the Social Media Mining for Health 2023 shared task. Our team addressed the first task, classifying tweets that self-report Covid-19 diagnosis. Our approach involves a classification model that incorporates diverse textual augmentations and utilizes R-drop to augment data and mitigate overfitting, boosting model efficacy. Our leading model, enhanced with R-drop and augmentations like synonym substitution, reserved words, and back translations, outperforms the task mean and median scores. Our system achieves an impressive F1 score of 0.877 on the test set.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- STonKGs: A Sophisticated Transformer Trained on Biomedical Text and Knowledge Graphs 95%
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 93%
- BERTMeSH: Deep Contextual Representation Learning for Large-scale High-performance MeSH Indexing with Full Text 92%
Similar papers in this journal
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 94%
- Use of large language models as a scalable approach to understanding public health discourse 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
Similar papers in this journal
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 94%
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 94%
- Accurate Prediction of Virus-Host Protein-Protein Interactions via a Siamese Neural Network Using Deep Protein Sequence Embeddings 90%
Similar papers in this journal
Similar papers in this journal
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 93%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 91%
- Robust Evaluation of Deep Learning-based Representation Methods for Survival and Gene Essentiality Prediction on Bulk RNA-seq Data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.