de novo variant calling identifies cancer mutation profiles in the 1000 Genomes Project
Ng, J. K.; Vats, P.; Fritz-Waters, E.; Padhi, E. M.; Payne, Z. L.; Leonard, S.; Sarkar, S.; West, M.; Prince, C.; Trani, L.; Jansen, M.; Vacek, G.; Samadi, M.; Harkins, T. T.; Pohl, C.; Turner, T. N.
Show abstract
Detection of de novo variants (DNVs) is critical for studies of disease-related variation and mutation rates. We developed a GPU-based workflow to rapidly call DNVs (HAT) and demonstrated its effectiveness by applying it to 4,216 Simons Simplex Collection (SSC) whole-genome sequenced parent-child trios from DNA derived from blood. In our SSC DNV data, we identified 78 {+/-} 15 DNVs per individual, 18% {+/-} 5% at CpG sites, 75% {+/-} 9% phased to the paternal chromosome of origin, and an average allele balance of 0.49. These calculations are all in line with DNV expectations. We sought to build a control DNV dataset by running HAT on 602 whole-genome sequenced parent-child trios from DNA derived from lymphoblastoid cell lines (LCLs) from the publicly available 1000 Genomes Project (1000G). In our 1000G DNV data, we identified 740 {+/-} 967 DNVs per individual, 14% {+/-} 4% at CpG sites, 61% {+/-} 11% phased to the paternal chromosome of origin, and an average allele balance of 0.41. Of the 602 trios, 80% had > 100 DNVs and we hypothesized the excess DNVs were cell line artifacts. Several lines of evidence in our data suggest that this is true and that 1000G does not appear to be a static reference. By mutation profile analysis, we tested whether these cell line artifacts were random and found that 40% of individuals in 1000G did not have random DNV profiles; rather they had DNV profiles matching B-cell lymphoma. Furthermore, we saw significant excess of protein-coding DNVs in 1000G in the gene IGLL5 that has already been implicated in this cancer. As a result of cell line artifacts, 1000G has variants present in DNA repair genes and at Clinvar pathogenic or likely-pathogenic sites. Our study elucidates important implications of the use of sequencing data from LCLs for both reference building projects as well as disease-related projects whereby these data are used in variant filtering steps.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Long-read whole-genome sequencing-based concurrent haplotyping and aneuploidy profiling of single cells 96%
- Genome-wide characterization of human minisatellite VNTRs: population-specific alleles and gene expression differences 96%
- Discovery of A Polymorphic Gene Fusion via Bottom-Up Chimeric RNA Prediction 96%
Similar papers in this journal
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 95%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 94%
- Hidden genomic diversity of SARS-CoV-2: implications for qRT-PCR diagnostics and transmission 94%
Similar papers in this journal
- Virus-derived variation in diverse human genomes 96%
- Missense variants causing Wiedemann-Steiner syndrome preferentially occur in the KMT2A-CXXC domain and are accurately classified using AlphaFold2 94%
- GWAS and 3D chromatin mapping identifies multicancer risk genes associated with hormone-dependent cancers 93%
Similar papers in this journal
- Allele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs 94%
- Pre-processing of paleogenomes: Mitigating reference bias and postmortem damage in ancient genome data 94%
- STRling: a k-mer counting approach that detects short tandem repeat expansions at known and novel loci 93%
Similar papers in this journal
- Detection of homozygous and hemizygous partial exon deletions by whole-exome sequencing 95%
- svCapture: Efficient and specific detection of very low frequency structural variant junctions by error-minimized capture sequencing 94%
- STAMP: a multiplex sequencing method for simultaneous evaluation of mitochondrial DNA heteroplasmies and content 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.