Blended Length Genome Sequencing (blend-seq): Combining Short Reads with Low-Coverage Long Reads to Maximize Variant Discovery
Magner, R.; Cunial, F.; Basu, S.; Paulsen, R.; Saponas, S.; Shand, M.; Lennon, N.; Banks, E.
Show abstract
We introduce blend-seq, a workflow for combining data from traditional short-read sequencing pipelines with low-coverage long reads, to improve variant discovery for single samples without the full cost of high-coverage long reads. We demonstrate that with only 4x long-read coverage augmenting 30x short reads, we can improve SNP discovery across the genome, exceeding performance beyond even high-coverage short reads (60x). For genotype-agnostic discovery of structural variants, we see a threefold improvement in recall while maintaining precision by using the low-coverage long reads on their own, and show how we can improve genotyping accuracy by adding in the short-read data. In addition, we demonstrate how the long reads can better phase these variants, incorporating long-context information in the genome to substantially outperform phasing with short reads alone. Our experiments highlight the complementary nature of short- and long-read technologies: the former contributing higher depth for genotyping and the latter better resolution of larger events or those in difficult regions.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dysgu: efficient structural variant calling using short or long reads 98%
- Enhanced Detection and Genotyping of Disease-Associated Tandem Repeats Using HMMSTR and Targeted Long-Read Sequencing 95%
- PCLIPtools: A Robust Framework for Identifying RNA-Protein Interaction Sites from PAR-CLIP experiments. 95%
Similar papers in this journal
Similar papers in this journal
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 95%
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 95%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.