Advanced Whole Genome Sequencing Using a Complete PCR-free Massively Parallel Sequencing (MPS) Workflow
Xia, Z.; Jiang, Y.; Drmanac, R.; Shen, H.; Liu, P.; Li, Z.; Chen, F.; Jiang, H.; Shi, S.; Xi, Y.; Li, Q.; Wang, X.; Zhao, J.; Liang, X.; Xie, Y.; Wang, L.; Tian, W.; Berntsen, T.; Luo, Y.; Gong, M.; Li, J.; Xu, C.; Dai, S.; Mi, Z.; Ren, H.; Lin, Z.; Chen, A.; Zhang, W.; Mu, F.; Xu, X.
Show abstract
BackgroundSystematic errors can be introduced from DNA amplification during massively parallel sequencing (MPS) library preparation and sequencing array formation. Polymerase chain reaction (PCR)-free genomic library preparation methods were previously shown to improve whole genome sequencing (WGS) quality on the Illumina platform, especially in calling insertions and deletions (InDels). We hypothesized that substantial InDel errors continue to be introduced by the remaining PCR step of DNA cluster generation. In addition to library preparation and sequencing, data analysis methods are also important for the accuracy of the output data.In recent years, several machine learning variant calling pipelines have emerged, which can correct the systematic errors from MPS and improve the data performance of variant calling. ResultsHere, PCR-free libraries were sequenced on the PCR-free DNBSEQ arrays from MGI Tech Co., Ltd. (referred to as MGI) to accomplish the first true PCR-free WGS which the whole process is truly not only PCR-free during library preparation but also PCR-free during sequencing. We demonstrated that PCR-based WGS libraries have significantly (about 5 times) more InDel errors than PCR-free libraries.Furthermore, PCR-free WGS libraries sequenced on the PCR-free DNBSEQ platform have up to 55% less InDel errors compared to the NovaSeq platform, confirming that DNA clusters contain PCR-generated errors.In addition, low coverage bias and less than 1% read duplication rate was reproducibly obtained in DNBSEQ PCR-free using either ultrasonic or enzymatic DNA fragmentation MGI kits combined with MGISEQ-2000. Meanwhile, variant calling performance (single-nucleotide polymorphisms (SNPs) F-score>99.94%, InDels F-score>99.6%) exceeded widely accepted standards using machine learning (ML) methods (DeepVariant or DNAscope). ConclusionsEnabled by the new PCR-free library preparation kits, ultra high-thoughput PCR-free sequencers and ML-based variant calling, true PCR-free DNBSEQ WGS provides a powerful solution for improving WGS accuracy while reducing cost and analysis time, thus facilitating future precision medicine, cohort studies, and large population genome projects.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 95%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 95%
- Performance Comparison Of Agilent New SureSelect All Exon v8 Probes With v7 Probes For Exome Sequencing 95%
Similar papers in this journal
Similar papers in this journal
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing 98%
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 96%
- Disentangling primer interactions improves SARS-CoV-2 genome sequencing by the ARTIC Network's multiplex PCR 95%
Similar papers in this journal
- Performance analysis of conventional and AI-based variant callers using short and long reads 98%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 97%
- Detection and characterization of copy number variants based on whole-genome sequencing by DNBSEQ platforms 96%
Similar papers in this journal
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 97%
- eccDNA Atlas: a comprehensive resource of eccDNA catalog 95%
- Comparing full variation profile analysis with the conventional consensus method in SARS-CoV-2 phylogeny 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.