Back

Pathogen exposure misclassification can bias association signals in GWAS of infectious diseases when using population-based common controls

Duchen, D.; Vergara, C.; Thio, C. L.; Kundu, P.; Chatterjee, N.; Thomas, D. L.; Wojcik, G. L.; Duggal, P.

2022-07-17 genetic and genomic medicine
10.1101/2022.07.14.22276656 medRxiv
Show abstract

Genome-wide association studies (GWAS) have been performed to identify host genetic factors for a range of phenotypes, including for infectious diseases. The use of population-based common controls from biobanks and extensive consortiums is a valuable resource to increase sample sizes in the identification of associated loci with minimal additional expense. Non-differential misclassification of the outcome has been reported when the controls are not well-characterized, which often attenuates the true effect size. However, for infectious diseases the comparison of cases to population-based common controls regardless of pathogen exposure can also result in selection bias. Through simulated comparisons of pathogen exposed cases and population-based common controls, we demonstrate that not accounting for pathogen exposure can result in biased effect estimates and spurious genome-wide significant signals. Further, the observed association can be distorted depending upon strength of the association between a locus and pathogen exposure and the prevalence of pathogen exposure. We also used a real data example from the hepatitis C virus (HCV) genetic consortium comparing HCV spontaneous clearance to persistent infection with both well characterized controls, and population-based common controls from the UK Biobank. We find biased effect estimates for known HCV clearance-associated loci and potentially spurious HCV clearance-associations. These findings suggest that the choice of controls is especially important for infectious diseases or outcomes that are conditional upon environmental exposures.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.