Back

KCFtools: Rapid alignment-free method for introgression screening and GWAS using k-mer profiles

Selvanayagam, S.; Quiroz-Chavez, J.; Ramirez Gonzalez, R. H.; Uauy, C.; Smit, S.; Schranz, M. E.

2025-11-03 bioinformatics
10.1101/2025.11.01.685998 bioRxiv
Show abstract

MotivationIn the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in non-overlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. ResultsWe systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional SNP-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with Genome-Wide Association Studies (GWAS) software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipelines potential for high-resolution, reference-agnostic population genetic analysis. Availabilityhttps://github.com/sivasubramanics/kcftools Contactc.s.sivasubramani@gmail.com

Published in Bioinformatics (predicted rank #2) · training set

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.