Back

Refining bias correction in genome-wide association analyses of case-control studies

Darbani, B.; Pedersen, O. B. V.; Ostrowski, S. R.; Tan, Q.; Andersen, V.

2025-09-04 bioinformatics
10.1101/2025.09.01.673522 bioRxiv
Show abstract

Genome-wide association studies are vulnerable to confounding factors. This study provides evidence-based guidance for minimizing bias associated with genetic relatedness, SNP-specific non-additive allelic interactions, predisposed genotypes among controls, and multi-allelic polymorphism in case-control studies. The analyses demonstrated that genetic similarity within case or control groups introduces experimental bias, whereas genetic relatedness across case-control samples reduces this bias. These findings establish a general framework that can filter genetically related sub-communities or paired samples, whilst preserving maximal statistical power. Moreover, the skewed odds ratios resulting from predisposed genotypes among controls underscored the importance of age-related filtering to minimize this confounding effect. To ensure accurate genetic estimates, such as polygenic risk scores, the identification of SNP-specific allelic interaction models was also emphasized in case-control studies, contingent on normalization for within-population differences in genotype frequencies. Finally, a strategy is recommended to capture genetic effects at multi-allelic genomic positions accurately.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.