Genomic sequence context differs between germline and somatic structural variants allowing for their differentiation in tumor samples without paired normals
Chukwu, W.; Lee, S.; Crane, A.; Zhang, S.; Mittra, I.; Imielinski, M.; Beroukhim, R.; Dubois, F.; Dalin, S.
Show abstract
Although several recent studies have characterized structural variants (SVs) in germline and cancer genomes, the features of SVs in these different contexts have not been directly compared. We examined similarities and differences between 2 million germline and 115 thousand tumor SVs from a cohort of 963 patients from The Cancer Genome Atlas (TCGA). We found significant differences in features related to their genomic sequences and localization that suggest differences between SV-generating processes and selective pressures. For example, we found that transposon-mediated processes shape germline much more than somatic SVs, while somatic SVs more frequently show features characteristic of chromoanagenesis. These differences were extensive enough to enable us to develop a classifier-"the great GaTSV"-that accurately distinguishes between germline and cancer SVs in tumor samples that lack a matched normal sample.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GASOLINE: detecting germline and somatic structural variants from long-reads data. 96%
- Short and long-read genome sequencing methodologies for somatic variant detection; genomic analysis of a patient with diffuse large B-cell lymphoma 96%
- Massively parallel identification of functionally consequential noncoding genetic variants in undiagnosed rare disease patients 92%
Similar papers in this journal
- svCapture: Efficient and specific detection of very low frequency structural variant junctions by error-minimized capture sequencing 94%
- Characterization and Mitigation of Fragmentation Enzyme-Induced Dual Stranded Artifacts 92%
- Detection of homozygous and hemizygous partial exon deletions by whole-exome sequencing 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.