Beyond reference bias: Making pangenomes accessible with PangyPlot
Mastromatteo, S.; Chirmade, S.; Roshandel, D.; Thiruvahindrapuram, B.; Wang, Z.; Patel, R. V.; Sung, W. W.; Hajianpour, A.; Wang, C.; Lin, F.; Keenan, K.; Avolio, J.; Eckford, P.; Ratjen, F.; Canadian Cystic Fibrosis Gene Modifier Consortium, ; Strug, L.
Show abstract
Linear reference genomes have standardized genomics research but remain limited by reference bias, which skews read mapping and variant discovery. This bias can distort the interpretation of genetic variation, particularly for populations that are genetically distant from the reference. Pangenome graphs, such as those generated by the Human Pangenome Reference Consortium (HPRC), mitigate this limitation by integrating diverse haplotypes into a unified representation of human genetic variation. However, the complexity of graph-based data and the lack of intuitive visualization tools have hindered broader adoption. Here we introduce PangyPlot, a genome browser that simplifies exploration of pangenome graphs by retaining linear-style navigation, integrating gene annotations, abstracting complex variation into interpretable structures, and employing a dynamic, physics-based layout optimization engine. We demonstrate its utility by constructing a chromosome 7 graph from 101 individuals with cystic fibrosis (CF), capturing a broad spectrum of genetic variation. Using PangyPlot, we visualized CF-relevant loci and compared results with existing graph viewers, highlighting its ability to display both base-level and large structural variation. With an additional 64 PacBio HiFi assemblies, we fine-mapped a repeat-dense CF modifier locus on chromosome 5, where PangyPlot was used in conjunction with graph-based analysis to identify a repeat expansion in the 5' end of EXOC3 that may promote G-quadruplex formation and affect gene expression. Together, these examples demonstrate PangyPlot s capacity to make populationlevel variation interpretable. To support broader use of graph-based resources, we also released a live public instance of PangyPlot preloaded with HPRC data (https://pangyplot.research.sickkids.ca/).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Inferring compound heterozygosity from large-scale exome sequencing data 94%
- Long read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits 94%
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 94%
Similar papers in this journal
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 94%
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 94%
- Multi-modal investigation of the schizophrenia-associated 3q29 genomic interval reveals global genetic diversity with unique haplotypes and segments that increase the risk for non-allelic homologous recombination 94%
Similar papers in this journal
- GREEN-DB: A framework for the annotation and prioritization of non-coding regulatory variants from whole-genome sequencing data 95%
- Long-read whole-genome sequencing-based concurrent haplotyping and aneuploidy profiling of single cells 95%
- Loss of critical developmental and human disease-causing genes in 58 mammals 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.