Back

Improving data archiving practices in ancient genomics

Bergström, A.

2023-05-16 genomics
10.1101/2023.05.15.540553 bioRxiv
Show abstract

The sequencing of ancient DNA from preserved biological remains is producing a rich record of past genetic diversity in humans and other species. However, unless the primary data is made available in public archives in an appropriate fashion, its long-term value will not be fully realised. I surveyed publicly archived data from 42 recent ancient genomics studies. I found that half of the studies archived incomplete subsets of the generated genomic data, preventing accurate replication and representing a loss of data of potential use for future research. None of the studies met all archiving criteria that could be considered best practice. Based on these results, I make six recommendations for data producers: 1) archive all sequencing reads, not just those that can be aligned to a reference genome, 2) archive read alignments as well, but as secondary analysis files linked to the underlying raw read files, 3) provide correct experiment metadata on how samples, libraries and sequencing runs relate to each other, 4) provide informative sample metadata in the public archives, 5) publish and archive data from screening, low-coverage, poorly performing and negative experiments, and 6) document data archiving choices in papers, and review these as part of peer review processes. Given the reliance on destructive sampling of finite material, I argue that ancient genomics studies have a particularly strong responsibility to ensure the longevity and reusability of generated data.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.