SMART: Statistical Mitogenome Assembly with Repeats
Alqahtani, F.; Mandoiu, I. I.
Show abstract
By using next-generation sequencing technologies it is possible to quickly and inexpensively generate large numbers of relatively short reads from both the nuclear and mitochondrial DNA contained in a biological sample. Unfortunately, assembling such whole-genome sequencing (WGS) data with standard de novo assemblers often fails to generate high quality mitochondrial genome sequences due to the large difference in copy number (and hence sequencing depth) between the mitochondrial and nuclear genomes. Assembly of complete mitochondrial genome sequences is further complicated by the fact that many de novo assemblers are not designed for circular genomes, and by the presence of repeats in the mitochondrial genomes of some species.\n\nIn this paper we describe the Statistical Mitogenome Assembly with Repeats (SMART) pipeline for automated assembly of complete circular mitochondrial genomes from WGS data. SMART uses an efficient coverage-based filter to first select a subset of reads enriched in mtDNA sequences. Contigs produced by an initial assembly step are filtered using BLAST searches against a comprehensive mitochondrial genome database, and used as \"baits\" for an alignment-based filter that produces the set of reads used in a second de novo assembly and scaffolding step. In the presence of repeats, the possible paths through the assembly graph are evaluated using a maximum-likelihood model. Additionally, the assembly process is repeated a user-specified number of times on re-sampled subsets of reads to select for annotation the reconstructed sequences with highest bootstrap support.\n\nExperiments on WGS datasets from a variety of species show that the SMART pipeline produces complete circular mitochondrial genome sequences with a higher success rate than current state-of-the art tools, even from low coverage WGS data. The pipeline is available through an easy-to-use web interface at https://neo.engr.uconn.edu/?tool_id=SMART.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MitoHiFi: a python pipeline for mitochondrial genome assembly from PacBio High Fidelity reads 97%
- GenErode: a bioinformatics pipeline to investigate genome erosion in endangered and extinct species 96%
- HapSolo: An optimization approach for removing secondary haplotigs during diploid genome assembly and scaffolding. 96%
Similar papers in this journal
- Significantly improving the quality of genome assemblies through curation 95%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
- Vulcan: Improved long-read mapping and structural variant calling via dual-mode alignment 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.