KILDA: identifying KIV-2 repeats from kmers
Molitor, C.; Labidi, T.; Rimbert, A.; Cariou, B.; Di Filippo, M.; Bardel, C.
Show abstract
MotivationHigh concentration of lipoprotein(a), a lipoprotein with proatherogenic properties, is an important risk factor for cardiovascular disease. This concentration is mostly genetically determined by a complex interplay between the number of Kringle-IV type 2 repeats and Lipoprotein(a)-affecting variants. Besides lipoprotein(a) plasma concentration, there is an unmet need to identify individuals most at risk based on their LPA genotype. ResultsWe developed KILDA, a Nextflow pipeline, to identify the number of Kringle-IV type 2 repeats and Lp(a)-affecting variants directly from kmers generated from FASTQ files. The pipeline was tested on the 1000 Genomes Project (n=2459) and results were equivalent to DRAGEN-LPA (R2=0.93). In-silico datasets proved the robustness of KILDAs predictions under different scenarios of sequencing coverage and quality. ConclusionKILDA is an open-source and free-to-use pipeline to identify the number of Kringle-IV type 2 repeats and lipoprotein(a)-associated variants. Its results are equivalent to DRAGEN-LPA, offering a free and robust tool for determining the LPA kringle number even when inputting low coverage libraries. AvailabilityKILDA is publicly available at https://github.com/HCL-HUBL/KILDA along with a recipe to build an Apptainer image containing all the required dependencies. Contact: corentin.molitor@chu-lyon.fr Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Identification of allele-specific KIV-2 repeats and impact on Lp(a) measurements for cardiovascular disease risk 93%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 92%
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 92%
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
- A Bioinformatics Pipeline for Estimating Mitochondria DNA Copy Number and Heteroplasmy Levels from Whole Genome Sequencing Data 94%
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 92%
Similar papers in this journal
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 93%
- Investigation of an LPA KIV-2 nonsense mutation in 11,000 individuals: the importance of linkage disequilibrium structure in LPA genetics. 91%
- REViewer: Haplotype-resolved visualization of read alignments in and around tandem repeats 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.