A robust benchmark for detecting low-frequency variants in the HG002 Genome In A Bottle NIST reference material.
Daniels, C.; Abdulkadir, A.; Cleveland, M. H.; McDaniel, J. H.; Jaspez, D.; Rubio-Rodriguez, L. A.; Munoz-Barrera, A.; Lorenzo-Salazar, J. M.; Flores, C.; Yoo, B.; Sahraeian, S. M. E.; Wang, Y.; Rossi, M.; Visvanath, A.; Murray, L.; Chen, W.-T.; Catreux, S.; Han, J.; Mehio, R.; Parnaby, G.; Carroll, A.; Chang, P.-C.; Shafin, K.; Cook, D. E.; Kolesnikov, A.; Brambrink, L.; Mootor, M. F. E.; Patel, Y.; Yamaguchi, T. N.; Boutros, P. C.; Sienkiewicz, K.; Foox, J.; Mason, C. E.; Lajoie, B.; Ruiz-Perez, C. A.; Kruglyak, S.; Zook, J. M.; Olson, N. D.
Show abstract
Somatic mosaicism is an important cause of disease, but mosaic and somatic variants are often challenging to detect because they exist in only a fraction of cells. To address the need for benchmarking subclonal variants in normal cell populations, we developed a benchmark containing mosaic variants in the Genome in a Bottle Consortium (GIAB) HG002 reference material DNA from a large batch of a normal lymphoblastoid cell line. First, we used a somatic variant caller with high coverage (300x) Illumina whole genome sequencing data from the Ashkenazi Jewish trio to detect variants in HG002 not detected in at least 5% of cells from the combined parental data. These candidate mosaic variants were subsequently evaluated using >100x BGI, Element, and PacBio HiFi data. High confidence candidate SNVs with variant allele fractions above 5% were included in the HG002 draft mosaic variant benchmark, with 13/85 occurring in medically relevant gene regions. We also delineated a 2.45 Gbp subset of the previously defined germline autosomal benchmark regions for HG002 in which no additional mosaic variants >2% exist, enabling robust assessment of false positives. The variant allele fraction of some mosaic variants is different between batches of cells, so using data from the homogeneous batch of reference material DNA is critical for benchmarking these variants. External validation of this mosaic benchmark showed it can be used to reliably identify both false negatives and false positives for a variety of technologies and detection algorithms, demonstrating its utility for optimization and validation. By adding our characterization of mosaic variants in this widely-used cell line, we support extensive benchmarking efforts using it in simulation, spike-in, and mixture studies.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 97%
- A Complete Pedigree-Based Graph Workflow for Rare Candidate Variant Analysis 96%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 96%
Similar papers in this journal
- Long-Read Structural and Epigenetic Profiling of a Kidney Tumor-Matched Sample with Nanopore Sequencing and Optical Genome Mapping 94%
- Lancet2: Improved and accelerated somatic variant calling with joint multi-sample local assembly graph 94%
- DoBSeqWF: A framework for sensitive detection of individual genetic variation in pooled sequencing data 94%
Similar papers in this journal
- Quartet DNA reference materials and datasets for comprehensively evaluating germline variants calling performance 95%
- Paragraph: A graph-based structural variant genotyper for short-read sequence data 95%
- TADA - a Machine Learning Tool for Functional Annotation based Prioritisation of Putative Pathogenic CNVs 95%
Similar papers in this journal
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 95%
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 94%
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.