Blended Genome Exome (BGE) as a Cost Efficient Alternative to Deep Whole Genomes or Arrays
DeFelice, M.; Grimsby, J.; Howrigan, D.; Yuan, K.; Chapman, S. B.; Stevens, C.; DeLuca, S.; Townsend, M.; Buxbaum, J.; Pericak-Vance, M. A.; Qin, S.; Stein, D. J.; Teferra, S.; Xavier, R.; Huang, H.; Martin, A.; Neale, B.
Show abstract
Genomic scientists have long been promised cheaper DNA sequencing, but deep whole genomes are still costly, especially when considered for large cohorts in population-level studies. More affordable options include microarrays + imputation, whole exome sequencing (WES), or low-pass whole genome sequencing (WGS) + imputation. WES + array + imputation has recently been shown to yield 99% of association signals detected by WGS. However, a method free from ascertainment biases of arrays or the need for merging different data types that still benefits from deeper exome coverage to enhance novel coding variant detection does not exist. We developed a new, combined, "Blended Genome Exome" (BGE) in which a whole genome library is generated, an aliquot of that genome is amplified by PCR, the exome regions are selected and enriched, and the genome and exome libraries are combined back into a single tube for sequencing (33% exome, 67% genome). This creates a single CRAM with a low-coverage whole genome (2-3x) combined with a higher coverage exome (30-40x). This BGE can be used for imputing common variants throughout the genome as well as for calling rare coding variants. We tested this new method and observed >99% r2 concordance between imputed BGE data and existing 30x WGS data for exome and genome variants. BGE can serve as a useful and cost-efficient alternative sequencing product for genomic researchers, requiring ten-fold less sequencing compared to 30x WGS without the need for complicated harmonization of array and sequencing data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Summix: A method for detecting and adjusting for population structure in genetic summary data 95%
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 95%
- A survey of rare epigenetic variation in 23,116 human genomes identifies disease-relevant epivariations and novel CGG expansions 95%
Similar papers in this journal
- Characterising tandem repeat complexities across long-read sequencing platforms with TREAT and otter 96%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 94%
- Prioritization of enhancer mutations by combining allele-specific chromatin accessibility with deep learning 94%
Similar papers in this journal
- Efficient phasing and imputation of low-coverage sequencing data using large reference panels 94%
- Long read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits 94%
- Inferring compound heterozygosity from large-scale exome sequencing data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.