Back

KILDA: identifying KIV-2 repeats from kmers

Molitor, C.; Labidi, T.; Rimbert, A.; Cariou, B.; Di Filippo, M.; Bardel, C.

2025-01-18 bioinformatics
10.1101/2025.01.17.631891 bioRxiv
Show abstract

MotivationHigh concentration of lipoprotein(a), a lipoprotein with proatherogenic properties, is an important risk factor for cardiovascular disease. This concentration is mostly genetically determined by a complex interplay between the number of Kringle-IV type 2 repeats and Lipoprotein(a)-affecting variants. Besides lipoprotein(a) plasma concentration, there is an unmet need to identify individuals most at risk based on their LPA genotype. ResultsWe developed KILDA, a Nextflow pipeline, to identify the number of Kringle-IV type 2 repeats and Lp(a)-affecting variants directly from kmers generated from FASTQ files. The pipeline was tested on the 1000 Genomes Project (n=2459) and results were equivalent to DRAGEN-LPA (R2=0.93). In-silico datasets proved the robustness of KILDAs predictions under different scenarios of sequencing coverage and quality. ConclusionKILDA is an open-source and free-to-use pipeline to identify the number of Kringle-IV type 2 repeats and lipoprotein(a)-associated variants. Its results are equivalent to DRAGEN-LPA, offering a free and robust tool for determining the LPA kringle number even when inputting low coverage libraries. AvailabilityKILDA is publicly available at https://github.com/HCL-HUBL/KILDA along with a recipe to build an Apptainer image containing all the required dependencies. Contact: corentin.molitor@chu-lyon.fr Supplementary informationSupplementary data are available at Bioinformatics online.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.