Widely used GWAS methods can be poorly suited to SNP-level localization under diffuse polygenic architecture in livestock
Wang, X.; Wang, J.; Tiezzi, F.; Huang, Y.; Huang, W.; Maltecca, C.; Jiang, J.
Show abstract
In livestock populations, genome-wide association studies (GWAS) can produce strong, apparently localized associations even when no truly discrete nearby causal effect exists. This occurs because small effective population sizes, strong family structure, long-range linkage disequilibrium (LD), and diffuse polygenic architecture can cause the effects of many variants to accumulate and be captured jointly across broad genomic intervals, making variant-level associations difficult to interpret biologically. Using real genotypes, we constructed a benchmark in which phenotypes were simulated under diffuse polygenic architecture across a genome partitioned into alternating effect and null windows, with central-null regions (at least 1 Mb away from effect-containing regions) positioned to detect long-range LD-driven signal propagation. We evaluated nine configurations of six GWAS methods (BOLT-LMM, REGENIE, fastGWA, FarmCPU, BLINK, and SLEMM) under this architecture. The central finding is that strong associations, of the kind normally read as evidence of nearby moderate- or large-effect variants, are produced by many of these methods even though the simulated signal is distributed across many tiny effects and cannot be localized to any single variant. The methods differed sharply in the extent of locus-level spillover: several produced large numbers of genome-wide significant loci within central-null regions, whereas the full-GRM mixed-model benchmark (SLEMM) suppressed this spillover almost entirely. These results show that, under a highly polygenic architecture with livestock-like LD, GWAS tool choice has major consequences for biological interpretation. When the goal is to localize biologically meaningful signals rather than to flag association peaks that may merely reflect tiny effects accumulated through LD across a broad block, methods that control long-range LD spillover should be prioritized.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Non-parametric polygenic risk prediction using partitioned GWAS summary statistics 97%
- Evaluating Multi-Ancestry Genome-Wide Association Methods: Statistical Power, Population Structure, and Practical Implications 97%
- mBAT-combo: a more powerful test to detect gene-trait associations from GWAS data 97%
Similar papers in this journal
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 97%
- Flashfm: A Flexible and Shared Information Fine-mapping Approach for Multiple Quantitative Traits 97%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 96%
Similar papers in this journal
- Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context 97%
- Eliciting priors and relaxing the single causal variant assumption in colocalisationanalyses 96%
- Enhancing Portability of Trans-Ancestral Polygenic Risk Scores through Tissue-Specific Functional Genomic Data Integration 96%
Similar papers in this journal
Similar papers in this journal
- Optimizing and benchmarking polygenic risk scores with GWAS summary statistics 96%
- Single locus theory of admixture is insufficient for the study of complex traits in admixed populations 96%
- HOPS: a quantitative score reveals pervasive horizontal pleiotropy in human genetic variation is driven by extreme polygenicity of human traits and diseases 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.