A pangenome-graph approach for mapping and imputing barley sequences
Sarria, J.; Amhal, H.; Ramirez, C. J.; Igartua, E.; Casas, A. M.; Contreras-Moreira, B.
Show abstract
Barley (Hordeum vulgare) is a key cereal crop with exceptional adaptation to diverse environments. With a large, highly repetitive diploid genome, barley presents challenges for pangenome representation. Starting from the reference genome MorexV3, we describe the construction of a barley graph (Pan20) representing the global diversity of landraces and cultivars captured in the public pangenome V1. For mapping arbitrary sequences, a greedy strategy is proposed that combines GMAP alignment followed by intersection with a Practical Haplotype Graph (PHG). This enables presence-absence variation detection and provides a consistent MorexV3 physical coordinate system across genotypes, enabling comparative analysis and visualization. For imputation of genomic data, the PHG approach relies on k-mer pseudo-alignment against the graph. Benchmarks show that Pan20 can accurately align barley genomic and transcriptomic sequences, including those not present in the Morex reference, revealing that a third of long genomic sequences map on non-reference genomes. Moreover, experiments with Genotyping by Sequencing and low-pass sequencing data indicate that FASTQ files can be efficiently mapped and imputed against the graph, preserving local haplotype context. This flexible and scalable graph framework allows barley researchers to explore genetic diversity beyond a single reference and facilitates analysis of diversity panels at the haplotype level, going beyond SNPs. Documentation and a Docker container are available at https://github.com/eead-csic-compbio/barleygraph. The graph sequence mapping utility was added to the Web application https://barleymap.eead.csic.es.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 94%
- AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data 93%
- A graph clustering algorithm for detection and genotyping of structural variants from long reads 93%
Similar papers in this journal
- The Personal Genome Project-UK: an open access resource of human multi-omics data 90%
- Construction, Deployment, and Usage of the Human Reference Atlas Knowledge Graph for Linked Open Data 90%
- Chromosome-scale genome assembly of the diploid oat Avena longiglumis reveals the landscape of repetitive sequences, genes and chromosome evolution in grasses 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.