Large Language Model-assisted text mining reveals bacterial pathogen diversity
Vos, M.; Goeker, M.; Bendall, R.; Costa, F.
Show abstract
Compiling and characterising the diversity of bacterial pathogens of humans is a critical challenge to tackle infection risk, especially in the context of global antimicrobial resistance, climate change, and changing demographics. Here, we present a scalable, automated pipeline that harnesses large language models (LLMs) to systematically mine the biomedical literature for information of human pathogenicity across the bacterial domain. By interrogating tens of thousands of PubMed abstracts we identify 1,222 species with at least one abstract documenting human infection, of which 783 species are supported by [≥]3 abstracts and are regarded as confirmed pathogens. We extract, summarise, visualise and interpret data on infection contexts using both expert-curated LLM prompts and unsupervised text vectorisation. We show that these methods enable fine-grained trait mapping across taxa, including quantifying the degree of specialism or generalism in body site specificity for different taxa and the classification of pathogen species into 75 pathogen types. An objective measure of the rate at which species are reported in the literature coupled to species clustering offers insights into the drivers of pathogen emergence. Our LLM-driven strategy generates an open, updatable, evidence-based catalogue of bacterial human pathogens and their ecological and clinical traits, providing a foundation for public health surveillance, diagnostics, and predictive modelling. This work demonstrates the potential of AI-assisted literature synthesis to transform our understanding of microbial diversity, including its impact on human health.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DRAM for distilling microbial metabolism to automate the curation of microbiome function 94%
- GIP: An open-source computational pipeline for mapping genomic instability from protists to cancer cells 94%
- Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity 93%
Similar papers in this journal
- Establishment and comparative genomics of a high-quality collection of mosquito-associated bacterial isolates -- MosAIC (Mosquito-Associated Isolate Collection) 94%
- Identifying and prioritizing potential human-infecting viruses from their genome sequences 90%
- iPHoP: an integrated machine-learning framework to maximize host prediction for metagenome-assembled virus genomes 90%
Similar papers in this journal
- Exploring SNP Filtering Strategies: The Influence of Strict vs Soft Core 94%
- Resolving plasmid-encoded carbapenem resistance dynamics and reservoirs in a hospital setting through nanopore sequencing 94%
- Exact mapping of Illumina blind spots in the Mycobacterium tuberculosis genome reveals platform-wide and workflow-specific biases 93%
Similar papers in this journal
Similar papers in this journal
- A global resource for genomic predictions of antimicrobial resistance and surveillance of Salmonella Typhi at Pathogenwatch 94%
- BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters 93%
- A convolutional neural network highlights mutations relevant to antimicrobial resistance in Mycobacterium tuberculosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.