Back

Linkage disequilibrium scaling improves robustness and power to detect genomic regions under selection.

Kemppainen, P. M.; Guillaume, F.

2026-01-21 evolutionary biology
10.64898/2026.01.19.700334 bioRxiv
Show abstract

Understanding the genetic basis of adaptive evolution is central to predicting how populations respond to environmental change. Genome scans and genotype-environment association methods are widely used to detect loci under selection, but because they often ignore correlations among loci (linkage disequilibrium, LD), their power decreases with increasing genetic structuring and demographic complexity. While LD contains valuable information about selection and can reduce dimensionality by summarizing signal across correlated markers, existing LD-based approaches are computationally demanding and rely on ad hoc threshold choices. Here, we introduce a novel LD-scaled association statistic, F ', which incorporates local LD directly into genotype-environment association tests using a fast permutation-based quantile transformation. Using forward-in-time simulations with known ground truth, we evaluate F ' across two widely used methods--latent factor mixed models (LFMM) and EMMAX--and assess performance at the level of outlier regions rather than individual SNPs. We further integrate uncertainty in parameter choice and inference method using a consistency-based framework. Across simulations, LD-scaling and joint inference substantially increased detection power and robust-ness, yielding up to an order-of-magnitude improvement under strong genetic structuring and nearly doubling performance on average across scenarios. Applying this framework to three- and nine-spined stickleback datasets, we recover well-established genomic regions associated with parallel marine-freshwater adaptation with markedly reduced background noise compared to standard analyses. Together, our results demonstrate that explicitly leveraging LD structure and integrating over analytical uncertainty provides a powerful and computationally efficient extension to genome scans for selection, improving robustness and interpretability in both simulated and empirical settings.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.