A novel algorithm to accurately classify metagenomic sequences
Saha, S.; Wang, Z.; Rajasekaran, S.
Show abstract
Widespread availability of next-generation sequencing (NGS) technologies has prompted a recent surge in interest in the microbiome. As a consequence, metagenomics is a fast growing field in bioinformatics and computational biology. An important problem in analyzing metagenomic sequenced data is to identify the microbes present in the sample and figure out their relative abundances. In this article we propose a highly efficient algorithm dubbed as "Hybrid Metagenomic Sequence Classifier" (HMSC) to accurately detect microbes and their relative abundances in a metagenomic sample. The algorithmic approach is fundamentally different from other state-of-the-art algorithms currently existing in this domain. HMSC judiciously exploits both alignment-free and alignment-based approaches to accurately characterize metagenomic sequenced data. To demonstrate the effectiveness of HMSC we used 8 metagenomic sequencing datasets (2 mock and 6 in silico bacterial communities) produced by 3 different sequencing technologies (e.g., HiSeq, MiSeq, and NovaSeq) with realistic error models and abundance distribution. Rigorous experimental evaluations show that HMSC is indeed an effective, scalable, and efficient algorithm compared to the other state-of-the-art methods in terms of accuracy, memory, and runtime. Availability of data and materialsThe implementations and the datasets we used are freely available for non-commercial purposes. They can be downloaded from: https://drive.google.com/drive/folders/132k5E5xqpkw7olFjzYwjWNjyHFrqJITe?usp=sharing
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large-scale Inference of Cell Lineage Trees and Genotype Calling from Noisy Single-Cell Data Using Efficient Local Search 97%
- Strobemers: an alternative to k-mers for sequence comparison 96%
- Sequence aligners can guarantee accuracy in almost O(m log n) time: a rigorous average-case analysis of the seed-chain-extend heuristic 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.