Strain-level sample characterisation using long reads and MAPQ scores
Hall, G. A.; Speed, T. P.; Woodruff, C. J.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWA simple but effective method for strain-level characterisation of microbial samples using long read data is presented. The method, which relies on having a non-redundant database of reference genomes, differentiates between strains within species and determines their relative abundance. It provides markedly better strain differentiation than that reported for the latest long read tools. Good estimates of relative abundances of highly similar strains present at less than 1% are achievable with as little as 1Gb of reads. Host contamination can be removed without great loss of sample characterisation performance. The method is simple and highly flexible, allowing it to be used for various different purposes, and as an extension of other characterisation tools. A code body implementing the underlying method is freely available.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Life at the extremes: Maximally divergent microbes with similar genomic signatures linked to extreme environments 96%
- Metagenomics-Toolkit: The Flexible and Efficient Cloud-Based Metagenomics Workflow featuring Machine Learning-Enabled Resource Allocation 96%
- ResistoXplorer: a web-based tool for visual, statistical and exploratory data analysis of resistome data 95%
Similar papers in this journal
- Functional Analysis of Metagenomes by Likelihood Inference (FAMLI) Successfully Compensates for Multi-Mapping Short Reads from Metagenomic Samples 97%
- Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets 96%
- cognac: rapid generation of concatenated gene alignments for phylogenetic inferencefrom large whole genome sequencing datasets 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.