FPfilter: A false-positive-specific filter for whole-genome sequencing variant calling from GATK
Tan, Y.; Zhang, Y.; Yang, H.; Yin, Z.
Show abstract
1.MotivationAs whole genome sequencing (WGS) is becoming cost-effective progressivelly, it has been applied increasingly in medical and scientific fields. Although the traditional variant-calling pipeline (BWA+GATK) has very high accuracy, false positives (FPs) are still an unavoidable problem that might lead to unfavorable outcomes, especially in clinical applications. As a result, filtering out FPs is recommended after variant calling. However, loss of true positives (TPs) is inevitable in FP-filtering methods, such as GATK hard filtering (GATK-HF). Therefore, to minimize the loss of TPs and maximize the filtration of FPs, a comprehensive understanding of the features of TPs and FPs, and building an improved model of classification are necessary. To obtain information about TPs and FPs, we used Platinum Genome (PT) as the mutation reference and its 300x deep sequenced dataset NA12878 as the simulation template. Then random sampling across depth gradients from NA12878 was performed to study the depth effect. ResultsFPs among heterozygous mutations were found to have pattern distinct from that of FPs among homozygous mutations. FPfilter makes use of this model to filter out FPs specifically. We evaluated FPfilter on a training dataset with depth gradients from NA12878 and a test dataset from NA12877 and NA24385. Compared with GATK-HF, FPfilter showed a significantly higher FP/TP filtration ratio and F-measure score. Our results indicate that FPfilter provides an improved model for distinguishing FPs from TPs and filters FPs with high specificity. AvailabilityFPfilter is freely available for download on GitHub (https://github.com/yuxiangtan/FPfilter). Users can easily install it from anaconda.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance analysis of conventional and AI-based variant callers using short and long reads 98%
- Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP 97%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 97%
Similar papers in this journal
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 96%
- Vulcan: Improved long-read mapping and structural variant calling via dual-mode alignment 95%
- Comparative analysis of seven short-reads sequencing platforms using the Korean Reference Genome: MGI and Illumina sequencing benchmark for whole-genome sequencing 95%
Similar papers in this journal
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 96%
- Comparing full variation profile analysis with the conventional consensus method in SARS-CoV-2 phylogeny 94%
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 94%
Similar papers in this journal
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 96%
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing 96%
- BC-store: a program for mgiseq barcode sets analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.