Back

Large language models enable consensus-level interpretation in metagenomic diagnostics

Steinig, E.; Krysiak, M.; Deo, K.; Duncan, A.; Prestedge, J.; Barr, J.; Moselen, J.; Khan, S. F.; Fernando, J. A.; Savic, I.; Yellapu, B.; Aziz, A.; Wirth, W.; Parry, J.; McDonald, A.; Lim, C.; Trevor, S.; Aw-Yeong, B.; McCluskey, G.; Moso, M.; Chan, E.; La Vita, S. L.; Bryant, P. A.; Crowe, A.; Maalim, R.; Velasquez Reyes, D.; Graham, M.; Williams, E.; Kwong, J. C.; Woolstencroft, R.; Slavin, M.; Lim, L. L.; Coin, L. J. M.; Caly, L.; Bond, K.; Kok Lim, C.; Stinear, T. P.; Williamson, D. A.; Ramachandran, P. S.

2026-07-31 infectious diseases
10.64898/2026.07.29.26358751 medRxiv
Show abstract

Abstract Metagenomic sequencing can detect a broad range of pathogens, but interpreting which detections are clinically relevant requires expert adjudication that is difficult to scale and standardize. Here we present diagnostic classifiers that formalize expert adjudication by combining structured decision trees with large language model reasoning to assign diagnoses and select pathogen candidates. We first developed a short-read metagenomic assay for sterile-site specimens (cerebrospinal and ocular fluid) in the META-GP study (Victoria, Australia, 2024-2025) and evaluated classifiers on a validation dataset (n = 96; clinical samples, spike-ins and controls). Locally deployed, open-weight reasoning models (Qwen3) achieved diagnostic performance comparable to expert consensus, improving with clinical context (n = 79, above experimental limit-of-detection; without clinical notes, 94.4% sensitivity, 95.4% specificity; with clinical notes, 97.2% sensitivity, 100% specificity). Automated adjudication enabled systematic benchmarking of computational parameters and regression testing for pathogen detection tasks. In a heterogeneous development cohort (n = 78), reviewers and classifiers identified clinically significant pathogens missed during routine testing. By reproducing consensus detections without requiring a full review panel, diagnostic classifiers enable scalable, standardized metagenomic interpretation that complements expert adjudications.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.