Back

Linked machine learning classifiers improve species classification of fungi when using error-prone long-reads on extendedmetabarcodes

Eenjes, T.; Hu, Y.; Irinyi, L.; Hoang, M. T.; Smith, L.; Linde, C.; Stone, E.; Rathjen, J. P.; Mashford, B.; Schwessinger, B.

2021-05-01 bioinformatics
10.1101/2021.05.01.442223 bioRxiv
Show abstract

BackgroundThe increased usage of error-prone long-read sequencing for metabarcoding of fungi has not been matched with adequate public databases and concomitant analysis approaches. We address this gap and present a proof-of-concept study for classifying fungal taxa using linked machine learning classifiers. We demonstrate the capability of linked machine learning classifiers to accurately classify species and strains using real-world and simulated fungal ribosomal DNA datasets, including plant and human pathogens. We benchmark our new approach in comparison to current alignment and k-mer based methods based on synthetic mock communities. We also assess real world applications of species identification in complex unlabelled datasets. ResultsOur machine learning approach assigned individual nanopore long-read amplicon sequences to fungal species with high recall rates and low false positive rates. Importantly, our approach successfully distinguished between closely-related species and strains when individual read errors were higher than the genetic distance between individual taxa, which the alignment and k-mer methods could not do. The machine learning approach showed an ability to identify key species with high recall rates, even in complex samples of unknown species composition. ConclusionsA proof of concept machine learning approach using a tree-descent approach on a decision tree of classifiers can identify known taxa with high accuracy, and precisely detect known target species from complex samples with high recall rates. We propose this approach is suitable for detecting the known knowns of pathogens or invasive species in any environment of mostly unknown composition, including agriculture and wild ecosystems.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.