CNV-PG: a machine-learning framework for accurate copy number variation predicting and genotyping
Wang, T.; Sun, J.; Zhang, X.; Wang, W.-J.; Zhou, Q.
Show abstract
MotivationCopy-number variants (CNVs) are one of the major causes of genetic disorders. However, current methods for CNV calling have high false-positive rates and low concordance, and a few of them can accurately genotype CNVs. ResultsHere we propose CNV-PG (CNV Predicting and Genotyping), a machine-learning framework for accurately predicting and genotyping CNVs from paired-end sequencing data. CNV-PG can efficiently remove false positive CNVs from existing CNV discovery algorithms, and integrate CNVs from multiple CNV callers into a unified call set with high genotyping accuracy. AvailabilityCNV-PG is available at https://github.com/wonderful1/CNV-PG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detection and characterization of copy number variants based on whole-genome sequencing by DNBSEQ platforms 98%
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 96%
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 95%
Similar papers in this journal
- Characterizing sensitivity and coverage of clinical WGS as a diagnostic test for genetic disorders 97%
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 92%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 92%
Similar papers in this journal
- Overcoming the pitfalls of NGS-based molecular diagnosis of Shwachman-Diamond syndrome 95%
- Evaluating discordant somatic calls across mutation discovery approaches to minimize false negative drug-resistant findings 94%
- Third Generation Cytogenetic Analysis (TGCA): diagnostic application of long-read sequencing. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.