Clustering of plasmid genomes for genomic epidemiology by using rearrangement distances, with pling
Frolova, D.; Iqbal, Z.
Show abstract
Integration of plasmids into genomic epidemiology is challenging, because there are no clearly defined evolving-units (equivalent to species), and because plasmids appear to evolve as much by structural change (rearrangements, insertions and deletions) as by mutation (1). Further, plasmids transfer horizontally between bacterial hosts (2), and thus a model beyond just a phylogeny is needed to integrate their genetic information with that of their hosts. Pling (3) is a tool designed to measure a genetic distance between plasmids that is related to how they empirically appear to evolve, by measuring the distance between two plasmids as the minimum number of structural changes needed to change one plasmid into the other (ignoring SNP differences). Having done this, it constructs a relatedness network of the plasmids under study, and then clusters them into groups that are credibly recently related. We give here a protocol for running pling, and how we integrate its information with plasmid typing and SNP information. Together, these provide a system for deciding which plasmids are worth treating as "the same plasmid" for the purposes of epidemiology, quantifying their relatedness in terms of rearrangements and SNPs, and then seeing how they are distributed across the host phylogeny.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rapid, robust plasmid verification by de novo assembly of short sequencing reads 93%
- Unveiling the Microbial Realm with VEBA 2.0: A modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic, and viral multi-omics from either short- or long-read sequencing 93%
- iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.