SMeta, a binning tool using single-cell sequences to aid reconstructing metageome species accurately
Zhang, Y.; Cheng, M.; Ning, K.
Show abstract
Because of the large volume and complex structure of metagenomic data, traditional binning methods are often hard to classify microbial metagenomes effectively. To deal with these challenges, introducing longer and more accurate single-cell sequencing data is a possible solution. Inspired by the existing MetaBAT2 tool, this study develops a new vector-based binning algorithm, SMeta, which uses both metagenomic and single-cell sequencing data. SMeta is specifically designed for eukaryotic microbial metagenomes, with the long reads characteristic of single-cell data. By introducing the segment tree data structure, the algorithm aligns long single-cell sequences with short metagenomic sequences quickly. This approach allows for the use of reference genomes from genomic databases to replace single-cell data, which makes more precise identification and reconstruction possible for small genome fragments, which are typically overlooked by traditional methods. Also, it might provide a higher purity sequence set for subsequent assemblies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 97%
- Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life 97%
- Detecting genomic deletions from high-throughput sequence data with unsupervised learning 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Real-time resolution of short-read assembly graph using ONT long reads 96%
- Deep6mA: a deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species 96%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 95%
Similar papers in this journal
- MetaMLP: A fast word embedding based classifier to profile target gene databases in metagenomic samples 95%
- HiCzin: Normalizing metagenomic Hi-C data and detecting spurious contacts using zero-inflated negative binomial regression 94%
- A simple way to find related sequences with position-specific probabilities 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.