The Collaborative Cross Graphical Genome
Su, H.; Chen, Z.; Rao, J.; Najarian, M.; Shorter, J. R.; Pardo-Manuel de Villena, F.; McMillan, L.
Show abstract
The mouse reference is one of the most widely used and accurately assembled mammalian genomes, and is the foundation for a wide range of bioinformatics and genetics tools. However, it represents the genomic organization of a single inbred mouse strain. Recently, inexpensive and fast genome sequencing has enabled the assembly of other common mouse strains at a quality approaching that of the reference. However, using these alternative assemblies in standard genomics analysis pipelines presents significant challenges. It has been suggested that a pangenome reference assembly, which incorporates multiple genomes into a single representation, are the path forward, but there are few standards for, or instances of practical pangenome representations suitable for large eukaryotic genomes. We present a pragmatic graph-based pangenome representation as a genomic resource for the widely-used recombinant-inbred mouse genetic reference population known as the Collaborative Cross (CC) and its eight founder genomes. Our pangenome representation leverages existing standards for genomic sequence representations with backward-compatible extensions to describe graph topology and genome-specific annotations along paths. It packs 83 mouse genomes (8 founders + 75 CC strains) into a single graph representation that captures important notions relating genomes such as identity-by-descent and highly variable genomic regions. The introduction of special anchor nodes with sequence content provides a valid coordinate framework that divides large eukaryotic genomes into homologous segments and addresses most of the graph-based position reference issues. Parallel edges between anchors place variants within a context that facilitates orthogonal genome comparison and visualization. Furthermore, our graph structure allows annotations to be placed in multiple genomic contexts and simplifies their maintenance as the assembly improves. The CC reference pangenome provides an open framework for new tool chain development and analysis.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- CGC1, a new reference genome for Caenorhabditis elegans 96%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
- HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads 95%
Similar papers in this journal
- Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny 95%
- A Simple Deep Learning Approach for Detecting Duplications and Deletions in Next-Generation Sequencing Data 94%
- BlobToolKit Interactive quality assessment of genome assemblies 93%
Similar papers in this journal
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 94%
- ConsHMM Atlas: conservation state annotations for major genomes and human genetic variation 94%
- Intraspecific de novo gene birth revealed by presence absence variant genes in Caenorhabditis elegans 93%
Similar papers in this journal
- ConVarT: a search engine for matching human genetic variants with variants from non-human species 94%
- Local assembly of long reads enables phylogenomics of transposable elements in a polyploid cell line 94%
- Laboratory evolution of the bacterial genome structure through insertion sequence activation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.