Bacterial strain nomenclature in the genomic era: Life Identification Numbers using a gene-by-gene approach
Palma, F.; Hennart, M.; Jolley, K. A.; Crestani, C.; Wyres, K. L.; Bridel, S.; Yeats, C. A.; Brancotte, B.; Raffestin, B.; David, S.; Lam, M. M. C.; Izdebski, R.; Passet, V.; Rodrigues, C.; Rethoret-Pasty, M.; Maiden, M. C. J.; Aanensen, D. M.; Holt, K. E.; Criscuolo, A.; Brisse, S.
Show abstract
Unified strain taxonomies are needed for the epidemiological surveillance of bacterial pathogens and international communication in microbiological research. Core genome multilocus sequence typing (cgMLST) holds great promise for standardized high-resolution strain genotyping. However, this approach faces challenges including classification instability and disconnection of new nomenclature from widely adopted classical MLST identifiers. This essay discusses the cgMLST-based Life Identification Number (LIN) method, recently proposed as a stable multilevel strain taxonomy system applicable to most bacterial pathogens. We describe how LIN codes are implemented and used in practice for precise strain definitions and epidemiological tracking. Glossary Multilocus sequence typing (MLST)A genotyping method applied mostly to microbial strains to study population structure and epidemiology, based on comparing the nucleotide sequences of a small number (typically seven) of housekeeping protein-coding genes. In MLST, allele numbers are assigned to each sequence variant (allele) of a given gene. The MLST genotype of a bacterial strain is defined by the combination of the allele numbers observed at the genes that are included in the genotyping scheme. A sequence type (ST) is assigned to each unique combination of alleles, called an MLST profile. MLST was invented in 1998 and became a de-facto standard taxonomy of bacterial strains, albeit at low resolution. Core genome MLSTAn extension of MLST that analyzes sequence variation across hundreds to thousands of conserved (core) genes, shared by all strains of a species, providing higher resolution typing for genomic epidemiology and evolutionary studies. cgMLST schemes typically comprise 2000 to 4000 genes, depending on the genome size and genetic variation (in terms of presence/absence of genes) within bacterial species. A core genome sequence type (cgST) can be assigned to unique cgMLST profiles, i.e., a unique combination of cgMLST allelic numbers. Whole Genome Sequencing (WGS)A method that determines the complete DNA sequence of an organisms genome in a single process, providing comprehensive information for comparative genetic analyses based on cgMLST or other analytic methods. Single Nucleotide Polymorphisms (SNPs)Variations at a single base position in the DNA sequence among individuals isolates, strains or species, used as genetic markers for studying for example, evolutionary relationships or strain identity. Average nucleotide identity (ANI)A measure of genomic similarity between two organisms, calculated as the average percentage of identical nucleotides in orthologous genomic regions; commonly used to assess species-level relatedness in prokaryotes. TaxonomyHere, we apply the word taxonomy to bacterial strains as a system of classifying, naming and identifying strains based on shared genetic characteristics as defined by e.g., cgMLST.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- No Assembly Required: Using BTyper3 to Assess the Congruency of a Proposed Taxonomic Framework for the Bacillus cereus group with Historical Typing Methods 96%
- ARGContextProfiler: Extracting and Scoring the Genomic Contexts of Antibiotic Resistance Genes using Assembly Graphs 94%
- ProkBERT Family: Genomic Language Models for Microbiome Applications 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.