Back

FAIR enough? A perspective on the status of nucleotide sequence data and metadata on public archives

Hassenrueck, C.; Poprick, T.; Helfer, V.; Molari, M.; Meyer, R.; Kostadinov, I.

2021-09-24 molecular biology Community evaluation
10.1101/2021.09.23.461561 bioRxiv
Show abstract

Knowledge derived from nucleotide sequence data is increasing in importance in the life sciences, as well as decision making (mainly in biodiversity policy). Metadata standards have been established to facilitate sustainable sequence data management according to the FAIR principles (Findability, Accessibility, Interoperability, Reusability). Here, we review the status of metadata available for raw read Illumina amplicon and whole genome shotgun sequencing data derived from ecological metagenomic material that are accessible at the European Nucleotide Archive (ENA), as well as the compliance of the primary sequence data (fastq files) with data submission requirements. While overall basic metadata, such as geographic coordinates, were retrievable in 98% of the cases for this type of sequence data, interoperability was not always ensured and other (mainly conditionally) mandatory parameters were often not provided at all. Metadata standards, such as the Minimum Information about any(x) Sequence (MIxS), were only infrequently used despite a demonstrated positive impact on metadata quality. Furthermore, the sequence data itself did not meet the prescribed requirements in 31 out of 39 studies that were manually inspected. To tackle the most immediate needs to improve FAIR sequence data management, we provide a list of minimal suggestions to researchers, research institutions, funding agencies, reviewers, publishers, and databases, that we believe might have a potentially large positive impact on sequence data and metadata FAIRness, which is crucial for further research and its derived applications.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Molecular Ecology Resources
171 papers in training set
Top 0.1%
26.4%
2
Metabarcoding and Metagenomics
14 papers in training set
Top 0.1%
18.4%
3
Ecological Indicators
21 papers in training set
Top 0.1%
11.8%
50% of probability mass above
4
PeerJ
308 papers in training set
Top 0.8%
5.5%
5
PLOS ONE
5266 papers in training set
Top 36%
3.5%
6
Environmental DNA
56 papers in training set
Top 0.2%
3.2%
7
Water Research
79 papers in training set
Top 0.6%
2.0%
8
Methods in Ecology and Evolution
176 papers in training set
Top 1.0%
1.9%
9
Frontiers in Microbiology
427 papers in training set
Top 5%
1.7%
10
Scientific Reports
3612 papers in training set
Top 54%
1.7%
11
GigaScience
212 papers in training set
Top 3%
1.5%
12
Limnology and Oceanography: Methods
11 papers in training set
Top 0.1%
1.5%
13
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
14
Frontiers in Bioinformatics
49 papers in training set
Top 0.9%
1.1%
15
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
16
Frontiers in Marine Science
62 papers in training set
Top 0.8%
1.1%
17
Ecology and Evolution
267 papers in training set
Top 5%
1.0%
18
Environmental Microbiome
29 papers in training set
Top 0.7%
0.8%
19
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
20
Biology Methods and Protocols
61 papers in training set
Top 2%
0.8%
21
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
22
F1000Research
88 papers in training set
Top 5%
0.6%
23
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.6%
24
BMC Medical Research Methodology
47 papers in training set
Top 2%
0.6%