Somatic Truth Data from Cell Lineage
Shand, M.; Soto, J.; Lichtenstein, L.; Benjamin, D.; Farjoun, Y.; Brody, Y.; Maruvka, Y. E.; Blainey, P. C.; Banks, E.
Show abstract
Existing somatic benchmark datasets for human sequencing data use germline variants, synthetic methods, or expensive validations, none of which are satisfactory for providing a large collection of true somatic variation across a whole genome. Here we propose a dataset of short somatic mutations, that are validated using a known cell lineage. The dataset contains 56,974 (2,687 unique) Single Nucleotide Variations (SNV), 6,370 (316 unique) small Insertions and Deletions (Indels), and 144 (8 unique) Copy Number Variants (CNV) across 98 in silico mixed truth sets with a high confidence region covering 2.7 gigabases per mixture. The data is publicly available for use as a benchmarking dataset for somatic short mutation discovery pipelines.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BATCAVE: Calling somatic mutations with a tumor- and site-specific prior 94%
- A Bioinformatics Pipeline for Estimating Mitochondria DNA Copy Number and Heteroplasmy Levels from Whole Genome Sequencing Data 93%
- Characterization and Mitigation of Fragmentation Enzyme-Induced Dual Stranded Artifacts 93%
Similar papers in this journal
- annoFuse: an R Package to annotate, prioritize, and interactively explore putative oncogenic RNA fusions 94%
- Algorithmic improvements for discovery of germline copy number variants in next-generation sequencing data 94%
- wg-blimp: an end-to-end analysis pipeline for whole genome bisulfite sequencing data 94%
Similar papers in this journal
Similar papers in this journal
- CaBagE: a Cas9-based Background Elimination strategy for targeted, long-read DNA sequencing 94%
- cONcat: Computational reconstruction of concatenated fragments from long Oxford Nanopore reads 92%
- Different structural variant prediction tools yield considerably different results in Caenorhabditis elegans 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.