EukRibo: a manually curated eukaryotic 18S rDNA reference database to facilitate identification of new diversity
Berney, C.; Henry, N.; Mahe, F.; Richter, D. J.; de Vargas, C.
Show abstract
EukRibo is a manually curated, public reference database of small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes. EukRibo strives to include only sequences with verified, up-to-date taxonomic identifications, with a strong focus on protists, and relatively low genetic redundancy, to keep the database compact yet comprehensive. Environmental clone sequences representing previously identified novel diversity are accepted as reference sequences only if they have a precise lineage designation, useful for taxonomic annotation. EukRibo is part of a suite of public resources generated by the UniEuk project, which all follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo allows higher confidence in the taxonomic annotation of environmental metabarcodes, and should facilitate identification of new eukaryotic diversity at various taxonomic levels. The database is currently in version 2, and all versions are permanently stored and made available via the FAIR open platform Zenodo. It is our hope that EukRibo will help ongoing curation efforts of other 18S rDNA reference databases, and we welcome suggestions of corrections and new features to be included in subsequent versions.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- COInr and mkCOInr: Building and customizing a non-redundant barcoding reference database from BOLD and NCBI using a lightweight pipeline. 97%
- Validated removal of nuclear pseudogenes and sequencing artefacts from mitochondrial metabarcode data 96%
- debar, a sequence-by-sequence denoiser for COI-5P DNA barcode data 95%
Similar papers in this journal
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 93%
- metaVaR: introducing metavariant species models for reference-free metagenomic-based population genomics 92%
- Can we use it? On the utility of de novo and reference-based assembly of Nanopore data for plant plastome sequencing 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.