GVRP: Genome Variant Refinement Pipeline for variant analysis in non-human species using machine learning
Choi, J.; Zhou, B.; Song, G.
Show abstract
Many investigations of human disease require model systems such as non-human primates and their associated genome analyses. While DeepVariant excels in calling human genetic variations, its reliance on calibrating against known variants from previous population studies poses challenges for non-human species. To address this limitation, we introduce the Genome Variant Refinement Pipeline (GVRP), employing a machine learning-based approach to refine variant calls in non-human species. Rather than training separate variant callers for each species, we employ a machine learning model to accurately identify variations and filter out false positives from DeepVariant. In GVRP, we omit certain DeepVariant preprocessing steps and leverage the ground-truth Genome In A Bottle (GIAB) variant calls to train the machine learning model for non-human species genome variant refinement. We anticipate that GVRP will significantly expedite genome variation studies for non-human species.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 96%
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 95%
- DeepSSV: detecting somatic small variants in paired tumor and normal sequencing data with convolutional neural network 94%
Similar papers in this journal
- Vulcan: Improved long-read mapping and structural variant calling via dual-mode alignment 97%
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 96%
- Assessment of human diploid genome assembly with 10x Linked-Reads data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.