Overview of the 8th Social Media Mining for Health Applications (#SMM4H) Shared Tasks at the AMIA 2023 Annual Symposium
Klein, A. Z.; Banda, J. M.; Guo, Y.; Schmidt, A. L.; Xu, D.; Flores Amaro, J. I.; Rodriguez-Esteban, R.; Sarker, A.; Gonzalez-Hernandez, G.
Show abstract
The aim of the Social Media Mining for Health Applications (#SMM4H) shared tasks is to take a community-driven approach to address the natural language processing and machine learning challenges inherent to utilizing social media data for health informatics. The eighth iteration of the #SMM4H shared tasks was hosted at the AMIA 2023 Annual Symposium and consisted of five tasks that represented various social media platforms (Twitter and Reddit), languages (English and Spanish), methods (binary classification, multi-class classification, extraction, and normalization), and topics (COVID-19, therapies, social anxiety disorder, and adverse drug events). In total, 29 teams registered, representing 18 countries. In this paper, we present the annotated corpora, a technical summary of the systems, and the performance results. In general, the top-performing systems used deep neural network architectures based on pre-trained transformer models. In particular, the top-performing systems for the classification tasks were based on single models that were pre-trained on social media corpora. To facilitate future work, the datasets--a total of 61,353 posts--will remain available by request, and the CodaLab sites will remain active for a post-evaluation phase.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Medication information extraction using local large language models 96%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 94%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 94%
Similar papers in this journal
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 95%
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 94%
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 94%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 95%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 93%
Similar papers in this journal
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 96%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 94%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.