Back

Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads

Keskus, A.; Bryant, A.; Ahmad, T.; Yoo, B.; Aganezov, S.; Goretsky, A.; Donmez, A.; Lansdon, L. A.; Rodriguez, I.; Park, J.; Liu, Y.; Cui, X.; Gardner, J.; McNulty, B.; Sacco, S.; Shetty, J.; Zhao, Y.; Tran, B.; Narzisi, G.; Helland, A.; Cook, D. E.; Carroll, A.; Chang, P.-C.; Kolesnikov, A.; Molloy, E. K.; Pushel, I.; Guest, E.; Pastinen, T.; Shafin, K.; Miga, K. H.; Malikic, S.; Day, C.-P.; Robine, N.; Sahinalp, C.; Dean, M.; Farooqi, M. S.; Paten, B.; Kolmogorov, M.

2024-03-26 genetic and genomic medicine
10.1101/2024.03.22.24304756 medRxiv
Show abstract

Most current studies rely on short-read sequencing to detect somatic structural variation (SV) in cancer genomes. Long-read sequencing offers the advantage of better mappability and long-range phasing, which results in substantial improvements in germline SV detection. However, current long-read SV detection methods do not generalize well to the analysis of somatic SVs in tumor genomes with complex rearrangements, heterogeneity, and aneuploidy. Here, we present Severus: a method for the accurate detection of different types of somatic SVs using a phased breakpoint graph approach. To benchmark various short- and long-read SV detection methods, we sequenced five tumor/normal cell line pairs with Illumina, Nanopore, and PacBio sequencing platforms; on this benchmark Severus showed the highest F1 scores (harmonic mean of the precision and recall) as compared to long-read and short-read methods. We then applied Severus to three clinical cases of pediatric cancer, demonstrating concordance with known genetic findings as well as revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Published in Nature Biotechnology (predicted rank #18) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.