A new framework for detecting copy number variants from single nucleotide polymorphism data: 'rCNV', a versatile R package for paralogs and CNVs detection
Karunarathne, P.; Zhou, Q.; Schliep, K.; Milesi, P.
Show abstract
Studies show that copy number variants (CNVs), due to their ubiquitous presence in eukaryotes, contribute to phenotypic variation, environmental adaptation, and fuel species divergence at a previously unknown rate. However, the detection of CNVs in genomes, especially in non-model organisms is challenging due to the need for costly genomic resources and complex computational infrastructure. Therefore, to provide researchers with a low-cost and easily accessible resource, we developed a robust statistical framework and an R software package to detect CNVs using allelic-read depth from SNPs data. The core of the framework exploits i) the allelic-read depth ratio distribution in heterozygotes for individual SNPs and testing it against an expected distribution under a binomial sampling, and ii) SNPs showing an apparent excess of heterozygotes under Hardy-Weinberg equilibrium, to detect alleles in putatively multi-copy regions. The use of multiple statistical tests to find the deviation in allelic-read depth ratio distribution makes our method sensitive to sampling and aware of reference biases thereby minimizing false detection of CNVs. Our framework is well-catered for high throughput short-reads data, hence, most GBS technologies (e.g., RADseq, Exome-capture, WGS). As such, it allows calling CNVs from genomes of varying complexity. The framework is implemented in the R package "rCNV" which effortlessly automates the analysis. We trained our models on simulated data and tested on four datasets obtained from different sequencing technologies (i.e., RADseq: Chinook salmon - Oncorhynchus tshawytscha, American lobster - Homarus americanus, Exome-capture: Norway Spruce - Picea abies, and WGS: Malaria mosquito -Anopheles gambiae).
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RecView: an interactive R application for viewing and locating recombination positions using pedigree data 96%
- ARPEGGIO: Automated Reproducible Polyploid EpiGenetic GuIdance workflOw 94%
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 93%
Similar papers in this journal
- Robust and efficient software for reference-free genomic diversity analysis of GBS data on diploid and polyploid species 95%
- SambaR: an R package for fast, easy and reproducible population-genetic analyses of biallelic SNP datasets 94%
- Simulation with RADinitio Improves RADseq Experimental Design and Sheds Light on Sources of Missing Data 94%
Similar papers in this journal
- read_haps: using read haplotypes to detect same species contamination in DNA sequences. 95%
- unCOVERApp: an interactive graphical application for clinical assessment of sequence coverage at the base-pair level 95%
- PacRAT: A program to improve barcode-variant mapping from PacBio long reads using multiple sequence alignment 95%
Similar papers in this journal
- VarGenius-HZD allows accurate detection of rare homozygous or hemizygous deletions in targeted sequencing leveraging breadth of coverage 93%
- The FORCE panel: An all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications 90%
- CREPE (CREate Primers and Evaluate): a computational tool for large-scale primer design and specificity analysis 89%
Similar papers in this journal
- WeavePop: A bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations 95%
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 94%
- kGWASflow: a modular, flexible, and reproducible Snakemake workflow for k-mers-based GWAS 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.