De novo clustering of long-read amplicons improves phylogenetic insight into microbiome data
Hui, Y.; Nielsen, D. S.; Krych, L.
Show abstract
Long-read amplicon profiling through read classification limits phylogenetic analysis of amplicons while community analysis of multicopy genes, relying on unique molecular identifier (UMI) corrections, often demands deep sequencing. To address this, we present a long amplicon consensus analysis (LACA) workflow employing multiple de novo clustering approaches based on sequence dissimilarity. LACA controls the average error rate of corrected sequences below 1% for the Oxford Nanopore Technologies (ONT) R9.4.1 and ONT R10.3 data, 0.2% for ONT R10.4.1, and 0.1% for high-accuracy ONT Duplex and Pacific Biosciences (PacBio) circular consensus sequencing (CCS) data in both simulated 16S rRNA and real 16-23S rRNA amplicon datasets. In high-accuracy PacBio CCS data, the clustering-based correction matched UMI correction, while outperforming 4xUMI correction in noisy ONT R10.3 and R9.4.1 data. Notably, LACA preserved phylogenetic fidelity in long operational taxonomic units and enhanced microbiome-wide phenotype characterization for synthetic mock communities and human vaginal samples.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BinaRena: a dedicated interactive platform for human-guided exploration and binning of metagenomes 97%
- MetaPro: A scalable and reproducible data processing and analysis pipeline for metatranscriptomic investigation of microbial communities 96%
- Single Amplified Genome Catalog Reveals the Dynamics of Mobilome and Resistome in the Human Microbiome 95%
Similar papers in this journal
Similar papers in this journal
- Illumina Complete Long Read Assay yields contiguous bacterial genomes from human gut metagenomes 96%
- BiG-MAP: an automated pipeline to profile metabolic gene cluster abundance and expression in microbiomes 96%
- Longitudinal, Multi-platform Metagenomics Yields a High-quality Genomic Catalog and Guides an In Vitro Model for Cheese Communities 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.