Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly
Qin, Q.; Heinz, J.; Li, H.
Show abstract
Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV calling method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic SVs in normal samples and somatic SVs in cancer cell lines with little loss in sensitivity. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics. SignificanceWe introduced a novel long-read SV calling method that leverages pangenome and personal genome and greatly improves the accuracy of somatic SV calling.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- svCapture: Efficient and specific detection of very low frequency structural variant junctions by error-minimized capture sequencing 95%
- Lancet2: Improved and accelerated somatic variant calling with joint multi-sample local assembly graph 95%
- Needlestack: an ultra-sensitive variant caller for multi-sample next generation sequencing data 94%
Similar papers in this journal
Similar papers in this journal
- OctopusV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis 95%
- LongPhase: an ultra-fast chromosome-scale phasing algorithm for small and large variants 95%
- GGTyper: genotyping complex structural variants using short-read sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.