Graph construction method impacts variation representation and analyses in a bovine super-pangenome
Leonard, A. S.; Crysnanto, D.; Mapel, X. M.; Bhati, M.; Pausch, H.
Show abstract
Several models and algorithms have been proposed to build pangenomes from multiple input assemblies, but their impact on variant representation, and consequently downstream analyses, is largely unknown. We create multi-species "super-pangenomes" using pggb, cactus, and minigraph with the Bos taurus taurus reference sequence and eleven haplotype-resolved assemblies from taurine and indicine cattle, bison, yak, and gaur. We recover 221k nonredundant structural variations (SVs) from the pangenomes, of which 135k (61%) are common to all three. SVs derived from assembly-based calling show high agreement with the consensus calls from the pangenomes (96%), but validate only a small proportion of variations private to each graph. Pggb and cactus, which also incorporate base-level variation, have approximately 95% exact matches with assembly-derived small variant calls, which significantly improves the edit rate when realigning assemblies compared to minigraph. We use the three pangenomes to investigate 9,566 variable number tandem repeats (VNTRs), finding 63% have identical predicted repeat counts in the three graphs, while minigraph can over or underestimate the count given its approximate coordinate system. We examine a highly variable VNTR locus and show that repeat unit copy number impacts expression of proximal genes and non-coding RNA. Our findings indicate good consensus between the three pangenome methods but also show their individual strengths and weaknesses that need to be considered when analysing different types of variants from multiple input assemblies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery 97%
- Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel 95%
- Long-read sequencing and genome assembly of natural history collection samples and challenging specimens 95%
Similar papers in this journal
- Chromosome-length genome assembly and structural variations of the primal Basenji dog (Canis lupus familiaris) genome 96%
- Benchmarking long-read variant calling in diploid and polyploid genomes: insights from human and plants 94%
- Optimizing Cost-Effective Gene Expression Phenotyping Approaches in Cattle Using 3' mRNA Sequencing 94%
Similar papers in this journal
- Pangenome genotyped structural variation improves molecular phenotype mapping in cattle 96%
- Taurine pangenome uncovers a segmental duplication upstream of KIT associated with depigmentation in white-headed cattle 96%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 94%
Similar papers in this journal
Similar papers in this journal
- Accurate assembly of the olive baboon (Papio anubis) genome using long-read and Hi-C data 95%
- Near-chromosomal de novo assembly of Bengal tiger genome reveals genetic hallmarks of apex-predation 95%
- A new duck genome reveals conserved and convergently evolved chromosome architectures of birds and mammals 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.