Open-source solution for evaluation and benchmarking of large language models for public health
Espinosa, L.; Kebaili, D.; Consoli, S.; Kalimeri, K.; Mejova, Y.; Salathe, M.
Show abstract
Large language models (LLMs) have demonstrated remarkable capabilities in various natural language processing tasks, including text classification, information extraction, and sentiment analysis. However, most existing benchmarks are general-purpose and lack relevance for domain-specific applications such as public health. To address this gap, the objective of this study is to develop and test an open-source solution enabling intuitive, easy and rapid benchmarking of popular LLMs in public health contexts. An LLM prediction prototype was developed, supporting multiple popular LLMs. It enables users to upload datasets, apply prompts, and generate structured JSON outputs for LLM tasks. The application prototype built with R and the Shiny library facilitates automated LLM evaluation and benchmarking by computing key performance metrics for categorical and variables. We tested these prototypes on four public health use cases: stance detection towards vaccination in tweets and Facebook posts, detection of vaccine adverse reaction in BabyCenter forum posts, and extraction of epidemiological quantitative features from the World Health Organization Disease Outbreak News. Results revealed high variability in LLM performance depending on the task, dataset, and model, with no single LLM consistently outperforming others across all tasks. While larger models generally excelled, smaller models performed competitively in specific scenarios, highlighting the importance of task-specific model selection. This study contributes to the effective LLM integration in public health by providing a structured, user-friendly and scalable solution for LLM prediction, evaluation and benchmarking. Our findings underline the relevance of standardised and task-specific evaluation methods and model selection, and the use of clear and structured prompts to improve LLM performance in domain-specific use cases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 95%
- “Is this Herpes or Syphilis?”: Latent Dirichlet Allocation Analysis of Sexually Transmitted Disease-Related Reddit Posts During the COVID-19 Pandemic 94%
- Quantified Flu: an individual-centered approach to gaining sickness-related insights from wearable data 94%
Similar papers in this journal
- Use of large language models as a scalable approach to understanding public health discourse 97%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 93%
- COVID-19 Vaccination Data Management and Visualization Systems for Improved Decision-Making: Lessons Learnt from Africa CDC Saving Lives and Livelihoods Program 93%
Similar papers in this journal
- Towards a COVID-19 symptom triad: The importance of symptom constellations in the SARS-CoV-2 pandemic 94%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
- Using Twitter To Generate Signals For The Enhancement Of Syndromic Surveillance Systems: Semi-Supervised Classification For Relevance Filtering in Syndromic Surveillance 92%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 94%
- An automated Dashboard to improve laboratory COVID-19 diagnostics management 94%
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 91%
Similar papers in this journal
- Measuring the impact of nonpharmaceutical interventions on the SARS-CoV-2 pandemic at a city level: An agent-based computational modeling study of the City of Natal 92%
- A dynamic ensemble model for short-term forecasting in pandemic situations 92%
- The Role of Modelling and Analytics in South African COVID-19 Planning and Budgeting 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.