Back

Comparative analysis of single nucleotide polymorphisms and microsatellite markers for parentage verification and discovery within the equine Thoroughbred breed

Flynn, P.; Morrin-O'Donnell, R.; Weld, R.; Gargan, L. M.; Carlsson, J.; Daly, S.; Suren, H.; Siddavatam, P.; Gujjula, K. R.

2021-07-28 genetics
10.1101/2021.07.28.453868 bioRxiv
Show abstract

Short tandem repeat (STR), also known as microsatellite markers are currently used for genetic parentage verification within equine. Transitioning from STR to single nucleotide polymorphism (SNP) markers to perform equine parentage verification is now a potentially feasible prospect and a key area requiring evaluation is parentage testing accuracies when using SNP based methods, in comparison to STRs. To investigate, we utilised a targeted equine genotyping by sequencing (GBS) panel of 562 SNPs to SNP genotype 309 Thoroughbred horses - inclusive of 55 previously parentage verified offspring. Availability of STR profiles for all 309 horses, enabled comparison of parentage accuracies between SNP and STR panels. An average sample call rate of 97.2% was initially observed, and subsequent removal of underperforming SNPs realised a pruned final panel of 516 SNPs. Simulated trio and partial parentage scenarios were tested across 12-STR, 16-STR, 147-SNP and 516-SNP panels. False-positives (i.e. expected to fail parentage, but pass) ranged from 0% for 147-SNP and 516-SNP panels to 0.003% when using 12-STRs within trio parentage scenarios, and 0% for 516-SNPs to 1.6% for 12-STRs within partial parentage scenarios. Our study leverages targeted GBS methods to generate low-density equine SNP profiles and demonstrates the value of SNP based equine parentage analysis in comparison to STRs - particularly when performing partial parentage discovery.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.