An Extensive Sequence Dataset of Gold-Standard Samples for Benchmarking and Development
Baid, G.; Nattestad, M.; Kolesnikov, A.; Goel, S.; Yang, H.; Chang, P.-C.; Carroll, A.
Show abstract
Accurate standards and extensive development datasets are the foundation of technical progress. To facilitate benchmarking and development, we sequence 9 samples, covering the Genome in a Bottle truth sets on multiple instruments (NovaSeq, HiSeqX, HiSeq4000, PacBio Sequel II System) and sample preparations (PCR-Free, PCR-Positive) for both whole genome and multiple exome kits. We benchmark pipelines, quantifying strengths and limitations for sequencing and analysis methods. We identify variability within and between instruments, preparation methods, and analytical pipelines, across various sequencing depths. We discuss the relevance of this variability to downstream analyses, and strategies to reduce variability.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 96%
- Detecting Foldback Artifacts in Long-reads 96%
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 94%
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 95%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 95%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
Similar papers in this journal
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 95%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
- Illumina But With Nanopore: Sequencing Illumina libraries at high accuracy on the ONT MinION using R2C2 94%
Similar papers in this journal
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 95%
- A Bioinformatics Pipeline for Estimating Mitochondria DNA Copy Number and Heteroplasmy Levels from Whole Genome Sequencing Data 94%
- Lancet2: Improved and accelerated somatic variant calling with joint multi-sample local assembly graph 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.