Mining Social Media Data for Influenza Vaccine Effectiveness Using a Large Language Model and Chain-of-Thought Prompting
Xu, D.; Lopez-Garcia, G.; O'Connor, K.; Holston, H.; Klein, A. Z.; Flores Amaro, I.; Scotch, M.; Gonzalez-Hernandez, G.
Show abstract
Influenza vaccine effectiveness (VE) estimation plays a critical role in public health decision-making by quantifying the real-world impact of vaccination campaigns and guiding policy adjustments. Current approaches to VE estimation are constrained by limited population representation, selection bias, and delayed reporting. To address some of these gaps, we propose leveraging large language models (LLMs) with few-shot chain-of-thought (CoT) prompting to mine social media data for real-time influenza VE estimation. We annotated over 4,000 tweets from the 2020-2021 flu season using structured guidelines, achieving high inter-annotator agreement. Our best prompting strategy achieves F1 scores above 87% for identifying influenza vaccination status and test outcomes, outperforming traditional supervised fine-tuning methods by large margins. These findings indicate that LLM-based prompting approaches effectively identify relevant social media information for influenza VE estimation, offering a valuable real-time surveillance tool that complements traditional epidemiological methods.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 94%
- Information retrieval in an infodemic: the case of COVID-19 publications 94%
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 94%
Similar papers in this journal
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 96%
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 95%
- Structuring clinical text with AI: old vs. new natural language processing techniques evaluated on eight common cardiovascular diseases 92%
Similar papers in this journal
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 96%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 93%
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.