GBSmode: a pipeline for haplotype-aware analysis of genotyping-by-sequencing data
Yates, S. A.; Studer, B.
Show abstract
2AO_SCPLOWBSTRACTC_SCPLOW2.1 BO_SCPLOWACKGROUNDC_SCPLOWGenotyping-by-sequencing (GBS) has revolutionised molecular genetic analysis. It enables simultaneous genotyping of thousands of DNA markers in the genome of any species. In contrast to whole-genome shotgun sequencing, GBS exploits a restriction enzyme to reduce genome complexity and directs the sequencing to begin at fixed digestion sites. However, currently used tools for the analysis of GBS data, such as SAMtools, often neglect the fundamental technical differences between GBS and shotgun sequencing. 2.2 RO_SCPLOWESULTSC_SCPLOWHere we present GBSmode, a dedicated pipeline to call DNA sequence variants using whole-read information from GBS data. It removes false positives by incorporating biological features such as the ploidy level and the number of possible alleles in the population under investigation. Comparison of GBSmode with SAMtools in an F2 population of rice (Oryza sativa L.) showed both identified a similar number of polymorphisms (13,449 and 14,445, respectively) with a high overlap (8,143). However, differences were found in the number of read misalignments (8.0% and 14.3% for GBSmode and SAMtools, respectively) and genotyping errors (5.0% and 8.3% for GBSmode and SAMtools, respectively). Further tests in a bi-parental F1 population of cassava (Manihot esculenta Crantz) showed GBSmode found 31,489 polymorphic loci, whereas the number was higher with SAMtools (43,860). However, this difference was mainly attributable to GBSmode rejecting 11,695 loci that were biologically not possible. 2.3 CO_SCPLOWONCLUSIONSC_SCPLOWThis study shows that GBSmode is a versatile tool for the analysis of GBS data. Moreover, GBSmode was able to reduce genotyping errors arising from read misalignments by combining haplotype data with biological information. Whilst other tools may find more markers, GBSmode is designed for accuracy.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Robust and efficient software for reference-free genomic diversity analysis of GBS data on diploid and polyploid species 97%
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 95%
- SambaR: an R package for fast, easy and reproducible population-genetic analyses of biallelic SNP datasets 94%
Similar papers in this journal
- Versatile mapping-by-sequencing with Easymap v.2 96%
- A Partially Phase-Separated Genome Sequence Assembly of the Vitis Rootstock 'Börner' (Vitis riparia x Vitis cinerea) and its Exploitation for Marker Development and Targeted Mapping 96%
- Novel design of imputation-enabled SNP arrays for breeding and research applications supporting multi-species hybridisation 96%
Similar papers in this journal
- <underline>Genome-wide Imputation Using the Practical Haplotype Graph in the Heterozygous Crop Cassava</underline> 96%
- Sexual dimorphism and the effect of wild introgressions on recombination in cassava (Manihot esculenta Crantz) breeding germplasm 95%
- stuart: an R package for the curation of SNP genotypes from experimental crosses 95%
Similar papers in this journal
- A sorghum Practical Haplotype Graph facilitates genome-wide imputation and cost-effective genomic prediction 95%
- Imputation of Low-density Marker Chip Data in Plant Breeding: Evaluation of Methods Based on Sugar Beet 95%
- Potato dihaploids uncover diverse alleles to facilitate diploid potato breeding 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.