Back

PanScan: A Tool for Tertiary Analysis of Human Pangenome Graphs

Uddin, M.; Balan, B.; Hanif, M.; Hashmi, A.; Kumail, M.; Bineshaq, S.; Alyazeedi, T.; Naveed, R.; Murtaza, M.; Jamalalail, B.; Elsokary, H.; Alobathani, M.; Mohamed, N.; Kondaramage, D.; Advani, D.; Ibrahim, A.; Shiyas, S.; Sares, N.; Alkhnbashi, O.; Duplessis, S.; Almarri, M.; Nassir, N.; Alsheikhali, A.

2025-05-07 bioinformatics
10.1101/2025.05.01.651685 bioRxiv
Show abstract

The genomic representation of populations across the globe is critical to ensuring a comprehensive and equitable human reference. Constructing a pangenome graph reference for different populations is the best approach to addressing local genomic diversities. Although major initiatives across continents are underway to construct pangenome graph references, the field lacks the necessary toolsets for tertiary analysis to characterize telomere-to-telomere (T2T) assemblies and the complexity of haplotypes. PanScan is a bioinformatics software package developed for human pangenome tertiary analysis. It includes multiple modules designed to detect duplicated gene sets from T2T assemblies, identify novel variants and sequences, as well as detect and visualize complex genomic regions through pangenome graph haplotype loops. We have used multiple pangenomes across different populations to assess the tertiary analysis and their accuracy. The tool is designed to streamline tertiary analysis and is compatible with multiple pangenome graph construction algorithms. PanScan is freely available on GitHub (https://github.com/CATG-Github/panscan), where users can provide human pangenome assemblies or VCF files as inputs for automated analyses through command-line operations on Linux systems. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=151 SRC="FIGDIR/small/651685v1_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@16cc212org.highwire.dtl.DTLVardef@13961d0org.highwire.dtl.DTLVardef@44d194org.highwire.dtl.DTLVardef@1b72aa_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.