Back

Observational epidemiological studies can mitigate genetic confounding with the genetic relatedness matrix

Patel, R. A.; Schraiber, J. G.; Pennell, M.; Edge, M. D.

2025-11-28 genetics
10.1101/2025.11.24.690292 bioRxiv
Show abstract

Observational studies are commonly used in psychology and epidemiology to identify risk factors correlated with health outcomes. However, these studies are vulnerable to confounding when shared genetic variation influences both the putative risk factor and outcome. Researchers have historically controlled for this type of genetic confounding using polygenic scores, but these scores are often noisy and biased estimators of a traits genetic component. Here, we develop a method that leverages a genetic relatedness matrix (GRM) to control genetic confounding when testing for non-genetic risk factors. In simulations, we find that our method outperforms existing approaches, particularly at sample sizes that are large by the standards of much human research but smaller than datasets often used in human genetics. We also demonstrate that existing methods are susceptible to poor GWAS portability, whereas our method is inherently robust to such concerns. Finally, we apply our method to the UK Biobank to re-analyze social risk factors for health outcomes in previously understudied cohorts. Significance statementLarge observational datasets are often used in epidemiology or social sciences to identify risk factors associated with health outcomes. However, these studies can be misleading when the putative risk factor and outcome are both complex genetic traits: a correlated genetic basis can create spurious associations between a putative risk factor and outcome. Existing approaches to address this problem typically rely on GWAS summary statistics, which require biobank-scale datasets. In this work, we introduce an alternative method that uses a genetic relatedness matrix (GRM) to control for genetic confounding directly. We show that our approach is closely related to existing methods but outperforms them in smaller datasets, making it an especially valuable tool for researchers who work on understudied traits and populations.

Published in Proceedings of the National Academy of Sciences (predicted rank #9) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.