Back

NOISYmputer: genotype imputation in bi-parental populationsfor noisy low-coverage next-generation sequencing data

Lorieux, M.; Gkanogiannis, A.; Fragoso, C.; Rami, J.-F.

2019-06-03 bioinformatics
10.1101/658237 bioRxiv
Show abstract

MotivationLow-coverage next-generation sequencing (LC-NGS) methods can be used to genotype bi-parental populations. This approach allows the creation of highly saturated genetic maps at reasonable cost, precisely localized recombination breakpoints, and minimize mapping intervals for quantitative-trait locus analysis.\n\nThe main issues with these genotyping methods are (1) poor performance at heterozygous loci, (2) a high percentage of missing data, (3) local errors due to erroneous mapping of sequencing reads and reference genome mistakes, and (4) global, technical errors inherent to NGS itself.\n\nRecent methods like Tassel-FSFHap or LB-Impute are excellent at addressing issues 1 and 2, but nonetheless perform poorly when issues 3 and 4 are persistent in a dataset (i.e. \"noisy\" data). Here, we present an algorithm for imputation of LC-NGS data that eliminates the need of complex pre-filtering of noisy data, accurately types heterozygous chromosomic regions, corrects erroneous data, and imputes missing data. We compare its performance with Tassel-FSFHap, LB-Impute, and Genotype-Corrector using simulated data and three real datasets: a rice single seed descent (SSD) population genotyped by genotyping by sequencing (GBS) by whole genome sequencing (WGS), and a sorghum SSD population genotyped by GBS.\n\nAvailabilityNOISYmputer, a Microsoft Excel-Visual Basic for Applications program that implements the algorithm, is available at mapdisto.free.fr. It runs in Apple macOS and Microsoft Windows operating systems.\n\nSupplementary files: Download link

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.