SVants: A long-read based method for structural variation detection in bacterial genomes
Hanson, B. M.; Johnson, J. S.; Leopold, S. R.; Sodergren, E.; Weinstock, G. M.
Show abstract
MotivationMobile genetic elements (MGEs) are genetic material that can transfer between bacterial cells and move to new locations within a single bacterial genome. These elements range from several hundred to tens of thousands of bases, and are often bordered by repeat regions, which makes resolving these elements difficult with short-read sequencing data. The development and availability of long-read sequencing technologies has opened up new opportunities in the study of structural variation but there is a lack of bioinformatics tools designed to take advantage of these longer reads.\n\nResultsWe present an assembly-free method for identifying the location of these MGEs when compared to any reference genome (including draft genomes). Using an artificially constructed Escherichia coli genome containing single and tandem-repeats of a Tn9 transposon, we demonstrate the ability of SVants to accurately identify multiple insertion sites as well as count the number of repeats of this MGE. Additionally, we show that SVants accurately identifies the transposon of interest, Tn9, but does not erroneously identify existing IS1 regions present within the chromosome of the E. coli artificial reference.\n\nAvailability and ImplementationSVants is available as open-source software at https://github.com/EpiBlake/SVants
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mBARq: a versatile and user-friendly framework for the analysis of DNA barcodes from transposon insertion libraries, knockout mutants and isogenic strain populations 95%
- Density-based binning of gene clusters to infer function or evolutionary history using GeneGrouper 92%
- Capturing variation in metagenomic assembly graphs with MetaCortex 92%
Similar papers in this journal
- Integrated population clustering and genomic epidemiology with PopPIPE 94%
- Bakta: Rapid & standardized annotation of bacterial genomes via alignment-free sequence identification 93%
- Platon: identification and characterization of bacterial plasmid contigs in short-read draft assembliesexploiting protein-sequence-based replicon distribution scores 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.