Benchmarking Large Language Models for Pathogen-Disease Classification in Post-Acute Infection Syndromes
Khalid, S. M.; Woelker, T.; Molano, L.-A. G.; Graf, S.; Keller, A.
Show abstract
Post-Acute Infection Syndromes (PAIS) are medical conditions that persist following acute infections from pathogens such as SARS-CoV-2, Epstein-Barr virus, and Influenza virus. Despite growing global awareness of PAIS and the exponential increase in biomedical literature, only a small fraction of this literature pertains specifically to PAIS, making the identification of pathogen-disease associations within such a vast, heterogeneous, and unstructured corpus a significant challenge for researchers. This study evaluated the effectiveness of large language models (LLMs) in extracting these associations through a binary classification task using a curated dataset of 1,000 manually labeled PubMed abstracts. We benchmarked a wide range of open-source LLMs of varying sizes (4B-70B parameters), including generalist, reasoning, and biomedical-specific models. We also investigated the extent to which prompting strategies such as zero-shot, few-shot, and Chain of Thought (CoT) methods can improve classification performance. Our results indicate that model performance varied by size, architecture, and prompting strategy. Zero-shot prompting produced the most reliable results: Mistral-Small-Instruct-2409 and Nemotron-70B achieved balanced accuracy scores of 0.81 and 0.80, respectively, along with macro-F1 scores of up to 0.80, while maintaining minimal invalid outputs. While few-shot and CoT prompting often degraded performance in generalist models, reasoning models such as DeepSeek-R1-Distill-Llama-70B and QwQ-32B demonstrated improved accuracy and consistency when provided with additional context.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 95%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 95%
- Machine Learning Made Easy (MLme): A Comprehensive Toolkit for Machine Learning-Driven Data Analysis 93%
Similar papers in this journal
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 96%
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 95%
- Relation extraction between bacteria and biotopes from biomedical texts with attention mechanisms and domain-specific contextual representations 95%
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 94%
- Robust Disease Prognosis via Diagnostic Knowledge Preservation: A Sequential Learning Approach 94%
- A cautionary tale about properly vetting datasets used in supervised learning predicting metabolic pathway involvement 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.