Back

Tier-based standards for FAIR sequence data and metadata sharing in microbiome research

Kim, L.; Lavrinienko, A.; Sebechlebska, Z.; Stoltenberg, S.; Bokulich, N.

2025-02-08 microbiology
10.1101/2025.02.06.636914 bioRxiv
Show abstract

Microbiome research is a growing, data-driven field within the life sciences. While policies exist for sharing microbiome sequence data and using standardized metadata schemes, compliance among researchers varies. To promote open research data best practices in microbiome research and adjacent communities, we (1) propose two tiered badge systems to evaluate data/metadata sharing compliance, and (2) developed an automated evaluation tool to determine adherence to data reporting standards in publications with amplicon and metagenome sequence data. In a systematic evaluation of publications (n = 2929) spanning human gut microbiome research, and in three case studies of soil and gut microbiota used to manually validate the evaluation tool (n = 370), we found nearly half of publications do not meet minimum standards for sequence data availability. Moreover, poor standardization of metadata creates a high barrier to harmonization and cross-study comparison. Using this badge system and evaluation tool, our proof-of-concept work exposes the (i) ineffectiveness of sequence data availability statements, and (ii) lack of consistent metadata reports used for annotation of microbial data. We highlight the need for improved practices and infrastructure that reduce barriers to data submission and maximize reproducibility in microbiome research. We anticipate that our tiered badge framework will promote dialogue regarding data sharing practices and facilitate microbiome data reuse, supporting best practices that make microbiome data FAIR. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=86 SRC="FIGDIR/small/636914v3_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@aaa4e8org.highwire.dtl.DTLVardef@130a3b1org.highwire.dtl.DTLVardef@4acc8aorg.highwire.dtl.DTLVardef@ba8688_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in Nucleic Acids Research (predicted rank #2) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.