Inference of ploidy by leveraging read depth from amplicon sequencing
Delomas, T. A.; Willis, S. C.; Schreier, A.; Narum, S.
Show abstract
Variation in ploidy occurs naturally in select plant and animal species. Ploidy variation can also occur spontaneously or be induced during artificial propagation of fish and shellfish. Studying species and systems that have variable ploidy requires techniques to infer ploidy of individuals. Massively parallel sequencing of biallelic SNPs has been used to infer ploidy, but existing techniques have several drawbacks. These include being limited to only comparing a fixed number of ploidies (diploidy, triploidy, and tetraploidy) and requiring that heterozygous genotypes in an individual be identified prior to ploidy inference. We describe a method of inferring ploidy from sequencing of biallelic SNPs based on beta-binomial mixture models. This method is generalized to apply to any ploidy and does not require prior identification of heterozygous genotypes. We demonstrate efficacy of this method for comparing ancestral octoploidy, decaploidy, and dodecaploidy (tetraploidy, pentaploidy, and hexaploidy for the sequenced SNPs) in white sturgeon and diploidy and triploidy in Chinook salmon with amplicon sequencing (GT-seq) data. Results indicated that ploidy could be reliably estimated for individuals based on distinct distribution of log-likelihood ratios (LLR) for known ploidy samples of both species that were tested. Confidence in ploidy estimates increased with sequencing depth. We encourage users to explore the sequencing depths and LLR critical values that provide reliable estimates of ploidy for a given organism and set of SNPs. We expect that the R package provided will empower studies of genetic variation and inheritance in organisms that vary in ploidy naturally or as a result of artificial propagation practices.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic and machine learning-based screening of aquaculture associated introgression into at-risk wild North American Atlantic salmon (Salmo salar) populations. 96%
- A GT-seq panel for walleye (Sander vitreus) provides a generalized workflow for efficient development and implementation of amplicon panels in non-model organisms. 95%
- RADSex: a computational workflow to study sex determination using Restriction Site-Associated DNA Sequencing data 94%
Similar papers in this journal
- An amplicon panel for high-throughput and low-cost genotyping of Pacific oyster 95%
- The generation of the first chromosome-level de-novo genome assembly and the development and validation of a 50K SNP array for North American Atlantic salmon 95%
- A high-quality chromosome-level genome assembly of rohu carp, Labeo rohita, and its utilization in SNP-based exploration of gene flow and sex determination 94%
Similar papers in this journal
- A population-level statistic for assessing Mendelian behavior of genotyping-by-sequencing data from highly duplicated genomes 94%
- ADMIXPIPE: Population analyses in ADMIXTURE for non-model organisms 90%
- USAT: a Bioinformatic Toolkit to Facilitate Interpretation and Comparative Visualization of Tandem Repeat Sequences 90%
Similar papers in this journal
- poolHelper: an R package to help in designing Pool-Seq studies 95%
- dartR v2: an accessible genetic analysis platform for conservation, ecology, and agriculture 94%
- Population genomic SNPs from epigenetic RADs: gaining genetic and epigenetic data from a single established next-generation sequencing approach 93%
Similar papers in this journal
- Patterns of linkage disequilibrium reveal genome architecture in chum salmon 95%
- Genomic prediction of growth in a commercially, recreationally, and culturally important marine resource, the Australian snapper (Chrysophrys auratus) 94%
- Performing Parentage Analysis for Polysomic Inheritances Based on Allelic Phenotypes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.