Back

SNPLift: Fast and accurate conversion of genetic variant coordinates across genome assemblies

Normandeau, E.; de Ronne, M.; Torkamaneh, D.

2023-06-14 genomics
10.1101/2023.06.13.544861 bioRxiv
Show abstract

MotivationThe advent of high-throughput sequencing technologies and the availability of reference genomes have provided an unprecedented opportunity to discover and genotype millions of genetic variants in hundreds or even thousands of samples. Variant calling, the identification of genetic variants from raw sequencing data, is both time-consuming and computationally demanding. Currently, reference genomes are evolving very rapidly and new assembly versions come out more and more frequently. To take advantage of new or improved reference genomes, raw reads alignments, genotype calling, and filtration must typically all be redone. This is a costly and time consuming operation that is not always viable when projects are under time constraints. ResultsHere, we introduce SNPLift, a bioinformatic pipeline that can quickly transfer the coordinate of nucleotide variants (SNPs and Indels) between different versions of reference genomes. We tested SNPLift on nine SNP datasets in VCF format from different species (Homo sapiens, Arabidopsis thaliana, Coregonus clupeaformis, Medicato truncatula, Oriza sativa, Salvelinus namaycush, Solanum lycopersicum, Zea mays, and Glycine max). Depending on the species, we achieved accurate lifting of variants ranging from 92.92% to 99.69%. Importantly, SNPLift significantly reduces the computational resources and time required for variant analysis compared to performing a complete re-analysis using a new reference genome. SNPLift offers a fast and efficient solution to leverage the benefits of updated or improved reference genomes. Availability and implementationSNPLift is available at https://github.com/enormandeau/snplift with its documentation. It contains a script that runs an automated test on a small dataset, composed of 190,443 SNPs in chromosome 1 of Medicago truncatula. SNPLift uses only common tools that are easy to install and works under Linux and MacOS.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.