Back

Consistent typing of plasmids with the mge-cluster pipeline

Arredondo-Alonso, S.; Gladstone, R. A.; Pontinen, A. K.; Gama, J. A.; Schurch, A. C.; Lanza, V. F.; Johnsen, P. J.; Samuelsen, O.; Tonkin-Hill, G.; Corander, J.

2022-12-19 bioinformatics
10.1101/2022.12.16.520696 bioRxiv
Show abstract

Extrachromosomal elements of bacterial cells such as plasmids are notorious for their importance in evolution and adaptation to changing ecology. However, high-resolution population-wide analysis of plasmids has only become accessible recently with the advent of scalable long-read sequencing technology. Current typing methods for the classification of plasmids remain limited in their scope which motivated us to develop a computationally efficient approach to simultaneously recognize novel types and classify plasmids into previously identified groups. Our method can easily handle thousands of input sequences which are compressed using a unitig representation in a de Bruijn graph. We provide an intuitive visualization, classification and clustering scheme that users can explore interactively. This provides a framework that can be easily distributed and replicated, enabling a consistent labelling of plasmids across past, present, and future sequence collections. We illustrate the attractive features of our approach by the analysis of population-wide plasmid data from the opportunistic pathogen Escherichia coli and the distribution of the colistin resistance gene mcr-1.1 in the plasmid population.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.