Population-level structural variant characterization from pangenome graph
Wang, S.; Xu, T.; Zhang, P.; Ye, K.
Show abstract
Population-level structural variant (SV) profiling is crucial in the era of pangenomes. However, identifying SVs from genome assemblies and pangenome graphs remains a significant challenge. Here we present Swave, a sequence-to-image, deep-learning based method that accurately resolves both simple and complex SVs, along with their population characteristics, from assembly-derived pangenome graphs. Swave introduces projection waves to summarize the dotplot images that capture mapping patterns between reference and SV-indicating alleles in pangenome. These images are analyzed by a recurrent neural network to distinguishes true SV signals from background noise introduced by genomic repeats. Swave demonstrates superior performance in both SV type classification and genotyping compared to existing methods. When applied to a healthy cohort (n=334) and a rare-disease cohort (n=574), Swave reveals complex and polymorphic SV patterns across human populations and identifies potentially pathogenic SVs. These advancements will facilitate the creation of comprehensive population-level SV catalogs, deepening our understanding of SVs in genetic diversity and disease associations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 97%
- Genome-wide Association Study of Long COVID 96%
- Genome-scale quantification and prediction of pathogenic stop codon readthrough by small molecules 96%
Similar papers in this journal
- False gene and chromosome losses affected by assembly and sequence errors 96%
- A read count-based method to detect multiplets and their cellular origins from snATAC-seq data 96%
- scDALI: Modelling allelic heterogeneity of DNA accessibility in single-cells reveals context-specific genetic regulation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.