Taming the reference genome jungle: the refget sequence collection standard
Campbell, D. R.; Cezard, T.; Gundersen, S.; Yates, A. D.; Davies, R. M.; Marshall, J.; Park, S.-H.; Wagner, A. H.; Love, M. I.; Thomas, R.; Hofmann, O.; Sheffield, N. C.
Show abstract
Reference genomes are foundational to genomics but suffer from widespread ambiguity and incompatibility due to inconsistent naming, undocumented differences, and lack of formal mechanisms for comparison. To address this, we introduce the GA4GH refget Sequence Collections (seqcol) standard. Refget seqcol is a framework for unambiguous representation, retrieval, and comparison of sequence collections such as reference genomes and transcriptomes. The seqcol standard comprises four components: a structured data schema, a canonical encoding algorithm that produces content-based, globally unique identifiers, a retrieval API, and a comparison protocol. This standard enables precise identification of sequence collections, even across decentralized or private systems, and allows compatibility assessments beyond exact identity, such as order-relaxed matches or shared coordinate systems. We applied the refget seqcol standard to 60 human and 36 mouse reference genomes sourced from major providers. Using digest-based comparisons, we quantified levels of similarity across attributes including sequence names, lengths, coordinate systems, and actual sequence content. Our analysis revealed some consistent subsets of sequences or coordinate systems, as well as substantial incompatibility among references and duplicate references under different names. To support adoption of refget seqcol, we provide a Python package implementing the full standard, a web API, and a comparison interface allowing users to assess local references against a curated database. This work offers a scalable, reproducible solution to the reference genome compatibility crisis, enabling improved transparency, reuse, and integration in genomic analyses. Refget seqcol enhances interoperability across tools and datasets, making genomic research more robust and reproducible.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Taxor: Fast and space-efficient taxonomic classification of long reads with hierarchicalinterleaved XOR filters 96%
- HISAT-3N: a rapid and accurate three-nucleotide sequence aligner 95%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
Similar papers in this journal
Similar papers in this journal
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 96%
- PEPhub: a database, web interface, and API for editing, sharing, and validating biological sample metadata 95%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.