Back

Pandemic-Potential Viruses are a Blind Spot for Frontier Open-Source LLMs

Luebbert, L.; Ektefaie, Y.; Rao, A. S.; Wilkason, C.; Nosamiefan, D.; Achonduh-Atijegbe, O.; Soumare, H.; Adebayo, A. P.; Olulaja, O.; Amadi, J.; Oyejide, N.; Olayiwola, F.; Henshaw, E.; Okocha, Y.; Nwachukwu, N.; Ewah, E. F.; Okoro, S.; Nwakpakpa, E.; Okokhere, P.; Iraoyah, K.; Okoeguale, J.; Dada, I.; Burris, A.; Zhao, K.; Laning, E.; van Amburg, C.; Cronan, P.; Fry, B.; Happi, C.; Ozonoff, A.; Sabeti, P. C.

2025-12-05 infectious diseases
10.64898/2025.12.04.25341642 medRxiv
Show abstract

We study large language models (LLMs) for front-line, pre-diagnostic infectious-disease triage, a critically understudied stage in clinical interventions, public health, and biothreat containment. We focus specifically on the operational decision of classifying symptomatic cases as viral vs. non-viral at first clinical contact, a critical decision point for resource allocation, quarantine strategy, and antibiotic use. We create a benchmark dataset of first-encounter cases in collaboration with multiple healthcare clinics in Nigeria, capturing high-risk viral presentations in low-resource settings with limited data. Our evaluations across frontier open-source LLMs reveal that (1) LLMs underperform standard tabular models and (2) case summaries and Retrieval Augmented Generation yield only modest gains, suggesting that naive information enrichment is insufficient in this setting. To address this, we demonstrate how models aligned with Group Relative Policy Optimization and a triage-oriented reward consistently improve baseline performance. Our results highlight persistent failure modes of general-purpose LLMs in pre-diagnostic triage and demonstrate how targeted reward-based alignment can help close this gap.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.