Back

Ultra-fast joint-genotyping with SparkGOR

Gudbjartsson, H.; Isleifsson, H. b.; Ragnarsson, B.; Guimaraes, R.; Wu, H.; Olafsdottir, H.; Stefasson, S. K.

2022-10-26 bioinformatics
10.1101/2022.10.25.513331 bioRxiv
Show abstract

MotivationOur aim was to simplify and speedup joint-genotyping, from sequence based variation data of individual samples, while maintaining as high sensitivity and specificity as possible. ResultsWe have leveraged versatile GOR data structures to store biallelic representations of variants and sequence read coverage in a very efficient way, allowing for very fast joint-genotyping that is an order of magnitude faster than any joint-genotyping method published to date. Furthermore, it can be easily extended and executed much faster in an incremental fashion. Concordance analysis based on the Genome In A Bottle (GIAB) samples shows favorable results when compared with the de-facto standard approach, using gVCF files and GATK joint-calling. Additionally, we have developed variant quality classification using XGBoost and variant training sets derived from the GIAB samples. The entire business logic is implemented efficiently and concisely in SparkGOR. AvailabilitySparkGOR is open-source and freely available at https://github.com/gorpipe. Contacthakon@genuitysci.com

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.