PanSVmerger: a flexible pipeline for merging multiallelic structural variants in pangenome graphs
Yang, T.; Shi, J.; Chen, Q.; Wu, D.; Tan, X.; Ruan, J.; Yang, C.
Show abstract
SummaryPangenome graphs capture extensive genetic diversity but introduce analytical challenges due to the redundant representation of structural variations (SVs). While existing tools effectively address cross-sample redundancy or cross-locus redundancy, none specifically target the intra-locus allelic redundancy inherent to pangenome graphs. Here, we present PanSVmerger, an open-source tool designed to consolidate redundant multiallelic SVs within individual loci using three complementary clustering strategies: adaptive k-mer-based Jaccard distance, global alignment distance via VSEARCH, and length distribution. Validation on HPRC pangenome data demonstrates that PanSVmerger effectively reduces multiallelic complexity (e.g., AC [≥] 3 loci from 62.4% to 4.7% using Strategy A) with a modest trade-off: Recall decreased from 97.13% to 93.58%, while precision improved from 94.95% to 96.56%, yielding an overall F1-score of 95.05%. These results demonstrate that PanSVmerger effectively consolidates redundant allele representations with only a minimal loss of sensitivity, making it well-suited for downstream applications that require clean, non-redundant variants. Availability and implementationPanSVmerger is implemented in Python 3.8+ and freely available under the MIT license at GitHub: https://github.com/tingting100/PanSVmerger. The software requires vcflib, bcftools, and optionally VSEARCH. Comprehensive documentation and tutorials are provided.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NucBreak: Location of structural errors in a genome assembly by using paired-end Illumina reads 93%
- DR2S: An Integrated Algorithm Providing Reference-Grade Haplotype Sequences from Heterozygous Samples 93%
- ILIAD: A suite of automated Snakemake workflows for processing genomic data for downstream applications 93%
Similar papers in this journal
- SatXplor - A comprehensive pipeline for satellite DNA analyses in complex genome assemblies 92%
- binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets 92%
- DeepSSV: detecting somatic small variants in paired tumor and normal sequencing data with convolutional neural network 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.