Linked machine learning classifiers improve species classification of fungi when using error-prone long-reads on extendedmetabarcodes
Eenjes, T.; Hu, Y.; Irinyi, L.; Hoang, M. T.; Smith, L.; Linde, C.; Stone, E.; Rathjen, J. P.; Mashford, B.; Schwessinger, B.
Show abstract
BackgroundThe increased usage of error-prone long-read sequencing for metabarcoding of fungi has not been matched with adequate public databases and concomitant analysis approaches. We address this gap and present a proof-of-concept study for classifying fungal taxa using linked machine learning classifiers. We demonstrate the capability of linked machine learning classifiers to accurately classify species and strains using real-world and simulated fungal ribosomal DNA datasets, including plant and human pathogens. We benchmark our new approach in comparison to current alignment and k-mer based methods based on synthetic mock communities. We also assess real world applications of species identification in complex unlabelled datasets. ResultsOur machine learning approach assigned individual nanopore long-read amplicon sequences to fungal species with high recall rates and low false positive rates. Importantly, our approach successfully distinguished between closely-related species and strains when individual read errors were higher than the genetic distance between individual taxa, which the alignment and k-mer methods could not do. The machine learning approach showed an ability to identify key species with high recall rates, even in complex samples of unknown species composition. ConclusionsA proof of concept machine learning approach using a tree-descent approach on a decision tree of classifiers can identify known taxa with high accuracy, and precisely detect known target species from complex samples with high recall rates. We propose this approach is suitable for detecting the known knowns of pathogens or invasive species in any environment of mostly unknown composition, including agriculture and wild ecosystems.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 94%
- Forage preference in two geographically co-occurring fungus gardening ants: a dietary DNA approach 94%
- Genome reconstruction of the non-culturable spinach downy mildew Peronospora effusa by metagenome filtering 94%
Similar papers in this journal
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 94%
- Evaluation of the accuracy of bacterial genome reconstruction with Oxford Nanopore R10.4.1 long-read-only sequencing 94%
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 94%
Similar papers in this journal
- dadasnake, a Snakemake implementation of DADA2 to process amplicon sequencing data for microbial ecology 96%
- IDseq - An Open Source Cloud-based Pipeline and Analysis Service for Metagenomic Pathogen Detection and Monitoring 95%
- PathoGFAIR: a collection of FAIR and adaptable (meta)genomics workflows for (foodborne) pathogens detection and tracking 94%
Similar papers in this journal
- Ribovore: ribosomal RNA sequence analysis for GenBank submissions and database curation 95%
- Profile hidden Markov model sequence analysis can help remove putative pseudogenes from DNA barcoding and metabarcoding datasets 94%
- Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets 94%
Similar papers in this journal
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 94%
- Genomic analysis of Coccomyxa viridis, a common low-abundance alga associated with lichen symbioses 92%
- Investigation of skin microbiota reveals Mycobacterium ulcerans-Aspergillus sp. trans-kingdom communication 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.