Pandemic-Potential Viruses are a Blind Spot for Frontier Open-Source LLMs
Luebbert, L.; Ektefaie, Y.; Rao, A. S.; Wilkason, C.; Nosamiefan, D.; Achonduh-Atijegbe, O.; Soumare, H.; Adebayo, A. P.; Olulaja, O.; Amadi, J.; Oyejide, N.; Olayiwola, F.; Henshaw, E.; Okocha, Y.; Nwachukwu, N.; Ewah, E. F.; Okoro, S.; Nwakpakpa, E.; Okokhere, P.; Iraoyah, K.; Okoeguale, J.; Dada, I.; Burris, A.; Zhao, K.; Laning, E.; van Amburg, C.; Cronan, P.; Fry, B.; Happi, C.; Ozonoff, A.; Sabeti, P. C.
Show abstract
We study large language models (LLMs) for front-line, pre-diagnostic infectious-disease triage, a critically understudied stage in clinical interventions, public health, and biothreat containment. We focus specifically on the operational decision of classifying symptomatic cases as viral vs. non-viral at first clinical contact, a critical decision point for resource allocation, quarantine strategy, and antibiotic use. We create a benchmark dataset of first-encounter cases in collaboration with multiple healthcare clinics in Nigeria, capturing high-risk viral presentations in low-resource settings with limited data. Our evaluations across frontier open-source LLMs reveal that (1) LLMs underperform standard tabular models and (2) case summaries and Retrieval Augmented Generation yield only modest gains, suggesting that naive information enrichment is insufficient in this setting. To address this, we demonstrate how models aligned with Group Relative Policy Optimization and a triage-oriented reward consistently improve baseline performance. Our results highlight persistent failure modes of general-purpose LLMs in pre-diagnostic triage and demonstrate how targeted reward-based alignment can help close this gap.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero Shot Health Trajectory Prediction Using Transformer 97%
- FedWeight: Mitigating Covariate Shift of Federated Learning on Electronic Health Records Data through Patients Re-weighting 96%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 93%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 93%
- Forecasting hospital-level COVID-19 admissions using real-time mobility data 92%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 97%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 92%
- Automated identification of abnormal infant movements from smart phone videos 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.