SVkhor: a unified framework for structural variant integration across long-read, short-read, and optical genome mapping data
Sharif Rahmani, E.; Thomas, Q.; Tisserant, E.; Vautrot, V.; Auclair, A.; Hounnondaho, F.-Z.; Castillon, E.; Faivre, L.; THAUVIN-ROBINET, C.; Vitobello, A.; Duffourd, Y.
Show abstract
Summary Multi-technology human genome structural variant (SV) discovery is challenged by differences in breakpoint resolution, allele representation, SV annotation, and VCF structure across various callers and platforms. Here, we present SVkhor, a software framework designed to merge outputs from multiple callers within each technology and integrate SV callsets across available short-read sequencing, long-read sequencing, and optical genome mapping data. SVkhor addresses these challenges through caller-aware normalization, within-technology merging, and cross-technology integration, producing compact, source-annotated SV catalogs suitable for benchmarking and downstream interpretation. Benchmarking using HG002 and analysis of a clinical trio demonstrate that SVkhor reduces redundant caller-level complexity while preserving technology-specific evidence, enabling the transition from heterogeneous SV callsets to interpretable sample- and family-level SV catalogs.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ClinSV: Clinical grade structural and copy number variant detection from whole genome sequencing data 95%
- Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches 95%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
Similar papers in this journal
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 96%
- MARRVEL-MCP enables natural language variant interpretation through autonomous workflow construction 95%
- MetaGLIMPSE: Meta Imputation of Low Coverage Sequencing Data for Modern and Ancient Genomes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.