Large language models enable consensus-level interpretation in metagenomic diagnostics
Steinig, E.; Krysiak, M.; Deo, K.; Duncan, A.; Prestedge, J.; Barr, J.; Moselen, J.; Khan, S. F.; Fernando, J. A.; Savic, I.; Yellapu, B.; Aziz, A.; Wirth, W.; Parry, J.; McDonald, A.; Lim, C.; Trevor, S.; Aw-Yeong, B.; McCluskey, G.; Moso, M.; Chan, E.; La Vita, S. L.; Bryant, P. A.; Crowe, A.; Maalim, R.; Velasquez Reyes, D.; Graham, M.; Williams, E.; Kwong, J. C.; Woolstencroft, R.; Slavin, M.; Lim, L. L.; Coin, L. J. M.; Caly, L.; Bond, K.; Kok Lim, C.; Stinear, T. P.; Williamson, D. A.; Ramachandran, P. S.
Show abstract
Abstract Metagenomic sequencing can detect a broad range of pathogens, but interpreting which detections are clinically relevant requires expert adjudication that is difficult to scale and standardize. Here we present diagnostic classifiers that formalize expert adjudication by combining structured decision trees with large language model reasoning to assign diagnoses and select pathogen candidates. We first developed a short-read metagenomic assay for sterile-site specimens (cerebrospinal and ocular fluid) in the META-GP study (Victoria, Australia, 2024-2025) and evaluated classifiers on a validation dataset (n = 96; clinical samples, spike-ins and controls). Locally deployed, open-weight reasoning models (Qwen3) achieved diagnostic performance comparable to expert consensus, improving with clinical context (n = 79, above experimental limit-of-detection; without clinical notes, 94.4% sensitivity, 95.4% specificity; with clinical notes, 97.2% sensitivity, 100% specificity). Automated adjudication enabled systematic benchmarking of computational parameters and regression testing for pathogen detection tasks. In a heterogeneous development cohort (n = 78), reviewers and classifiers identified clinically significant pathogens missed during routine testing. By reproducing consensus detections without requiring a full review panel, diagnostic classifiers enable scalable, standardized metagenomic interpretation that complements expert adjudications.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Convolutional neural networks quantify antibiotic resistance in Mycobacterium tuberculosis with diagnostic grade accuracy and predict treatment response 93%
- Identification of host-pathogen-disease relationships using a scalable Multiplex Serology platform in UK Biobank 92%
- Multiplexed detection of febrile infections using CARMEN 91%
Similar papers in this journal
- Phenotypic signatures in clinical data enable systematic identification of patients for genetic testing 92%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 92%
- Direct Antimicrobial Resistance Prediction from clinical MALDI-TOF mass spectra using Machine Learning 90%
Similar papers in this journal
Similar papers in this journal
- Evaluation of Oxford Nanopore Technologies workflows for genomic epidemiology of outbreak-associated bacterial isolates in the clinical setting 92%
- Tracking SARS-CoV-2 variants of concern in wastewater: an assessment of nine computational tools using simulated genomic data 92%
- Accelerating surveillance and research of antimicrobial resistance - an online repository for sharing of antimicrobial susceptibility data associated with whole genome sequences 92%
Similar papers in this journal
- High-Sensitivity Pan-Cancer AI Assessment of Lymph Node Metastasis via Uncertainty Quantification 94%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 93%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.