Tier-based standards for FAIR sequence data and metadata sharing in microbiome research
Kim, L.; Lavrinienko, A.; Sebechlebska, Z.; Stoltenberg, S.; Bokulich, N.
Show abstract
Microbiome research is a growing, data-driven field within the life sciences. While policies exist for sharing microbiome sequence data and using standardized metadata schemes, compliance among researchers varies. To promote open research data best practices in microbiome research and adjacent communities, we (1) propose two tiered badge systems to evaluate data/metadata sharing compliance, and (2) developed an automated evaluation tool to determine adherence to data reporting standards in publications with amplicon and metagenome sequence data. In a systematic evaluation of publications (n = 2929) spanning human gut microbiome research, and in three case studies of soil and gut microbiota used to manually validate the evaluation tool (n = 370), we found nearly half of publications do not meet minimum standards for sequence data availability. Moreover, poor standardization of metadata creates a high barrier to harmonization and cross-study comparison. Using this badge system and evaluation tool, our proof-of-concept work exposes the (i) ineffectiveness of sequence data availability statements, and (ii) lack of consistent metadata reports used for annotation of microbial data. We highlight the need for improved practices and infrastructure that reduce barriers to data submission and maximize reproducibility in microbiome research. We anticipate that our tiered badge framework will promote dialogue regarding data sharing practices and facilitate microbiome data reuse, supporting best practices that make microbiome data FAIR. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=86 SRC="FIGDIR/small/636914v3_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@aaa4e8org.highwire.dtl.DTLVardef@130a3b1org.highwire.dtl.DTLVardef@4acc8aorg.highwire.dtl.DTLVardef@ba8688_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 95%
- GSR-DB: a manually curated and optimised taxonomical database for 16S rRNA amplicon analysis 93%
- parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis 93%
Similar papers in this journal
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 93%
- Finding the right fit: A comprehensive evaluation of short-read and long-read sequencing approaches to maximize the utility of clinical microbiome data 93%
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 93%
Similar papers in this journal
- HVRLocator: A Computationally Efficient Tool for Identifying Hypervariable Regions in 16S rRNA Big Datasets 95%
- gNOMO2: a comprehensive and modular pipeline for integrated multi-omics analyses of microbiomes 94%
- IDseq - An Open Source Cloud-based Pipeline and Analysis Service for Metagenomic Pathogen Detection and Monitoring 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.