Back

KLinterSel: Intersection among candidates of different selective sweep detection methods

Carvajal-Rodriguez, A.; Rocha, S.; Pampin, M.; Martinez, P.; Caballero, A.

2025-12-01 genomics
10.1101/2025.08.21.671449 bioRxiv
Show abstract

Studies aiming to detect signals of selection in genomes often apply multiple methods to increase confidence in their results, typically selecting genomic regions that overlap across approaches. However, such overlap can be misleading when the genomic regions under study are not independent. In these cases, coincident candidates may arise from the structure of the data itself rather than from true methodological robustness. To address this issue, we present a statistical test that compares, for a given set of SNPs, the observed distance profile between candidate sites detected by different methods with the distance profile expected by chance for the same dataset. This test is implemented in the KLinterSel program, which additionally identifies clusters of sites jointly detected by several methods within a user-defined distance threshold. As a proof of concept, we applied KLinterSel to evaluate the overlap among candidates from four selection-detection methods investigating divergent selection associated with resistance to the parasite Marteilia cochillia in the common cockle (Cerastoderma edule). KLinterSel statistically evaluates and visualizes the agreement between observed and expected-by-chance distance profiles. It uses Pythons numerical libraries and vectorized operations for computational efficiency and includes multi-process parallelization options for memory-intensive datasets. Source code and documentation are available on GitHub (https://github.com/noosdev0/KLinterSel), and pre-built binaries for Windows, Linux, and macOS (arm64) facilitate broad accessibility.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.