LLMs for analyzing open text in global health surveys: why children are not accessing vaccine services in the DRC
Burstein, R. L.; Mufata, E.; Proctor, J. L.
Show abstract
This study evaluates the use of large language models (LLMs) to analyze free-text responses from large-scale global health surveys, using data from the Enquete de Couverture Vaccinale (ECV) household coverage surveys from 2020, 2021, 2022, and 2023 as a case study. We tested several LLM approaches varying from zero-shot and few-shot prompting, fine-tuning, and a natural language processing approach using semantic embeddings to analyze responses on reasons caregivers did not vaccinate their children. Performance ranged from 61.5% to 96% based on testing against a curated benchmarking dataset drawn from the ECV surveys, with accuracy improving when LLM models were fine-tuned or provided examples for few-shot learning. We show that even with as few as 20-100 examples, LLMs can achieve high accuracy in categorizing free-text responses. This approach offers significant opportunities for reanalyzing existing datasets and designing surveys with more open-ended questions, providing a scalable, cost-effective solution for global health organizations. Despite challenges with closed-source models and computational costs, the study underscores LLMs potential to enhance data analysis and inform global health policy.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Elucidating user behaviours in a digital health surveillance system to correct prevalence estimates 91%
- Foundation time series models for forecasting and policy evaluation in infectious disease epidemics 91%
- A prospective real-time transfer learning approach to estimate Influenza hospitalizations with limited data 91%
Similar papers in this journal
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 93%
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 93%
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 92%
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 91%
- Which COVID policies are most effective? A Bayesian analysis of COVID-19 by jurisdiction 91%
Similar papers in this journal
- Prioritizing countries for TB vaccine readiness research using a global stakeholder-centric approach 92%
- Identifying and preventing fraudulent responses in online public health surveys: Lessons learned during the COVID-19 pandemic 91%
- The Role of Modelling and Analytics in South African COVID-19 Planning and Budgeting 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.