KBeagle: An Adaptive Strategy and Tool for Improvement of Imputation Accuracy and Computing Efficiency
Qin, J.; Liu, X.; Liu, Y.; Wei, P.; Kangzhu, Y.; Zhong, J.; Wang, J.
Show abstract
With the development of molecular biology and genetics, deep sequencing technology has become the main way to discover genetic variation and reveal the molecular structure of genome. Due to the complexity of the whole genome segment structure, a large number of missing genotypes have appeared after sequencing, and these missing genotypes can be imputed by genotype imputation method. With the in-depth study of genotype imputation methods, computational intensive and computationally efficient imputation software come into being. Beagle software, as an efficient imputation software, is widely used because of its advantages of low memory consumption, fast running speed and relatively high imputation accuracy. K-Means clustering can divide individuals with similar population structure into a class, so that individuals in the same class can share longer haplotype fragments. Therefore, combining K-Means clustering algorithm with Beagle software can improve the interpolation accuracy. The Beagle and KBeagle method was used to compare the imputation efficiency. The KBeagle method presents a higher imputation matching rate and a shorter computing time. In the genome selection and heritability estimated section, the genotype dataset after imputed, unimputed, and with real genotype show similar prediction accuracy. However the estimated heritability using genotype dataset after imputed is closer to the estimation by the dataset with real genotype. We generated a compounds and efficient imputation method, which presents valuable resource for improvement of imputation accuracy and computing time. We envisage the application of KBeagle will be focus on the livestock sequencing study under strong genetic structure.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 95%
- Sample Size Impact (SaSii): an R script for estimating optimal sample sizes in population genetics and population genomics studies 95%
- Genome-wide DNA polymorphisms in four Actinidia arguta genotypes based on whole-genome re-sequencing 95%
Similar papers in this journal
- Multifactorial Methods Integrating Haplotype and Epistasis Effects for Genomic Estimation and Prediction of Quantitative Traits 94%
- Signature of selection in composite Vrindavani cattle of India 94%
- Improving short and long term genetic gain by accounting for within family variance in optimal cross selection 93%
Similar papers in this journal
- Analysis of the Genetic Structure and Diversity of Upland Cotton Groups in Different Planting Areas Based on SNP Markers 95%
- Analysis of selection signatures reveals important insights into the adaptability of high-altitude Indian sheep breed Changthangi 93%
- Systems Biology under heat stress in Indian Cattle 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.