Back

Identifying direct risk factors in UK Biobank with simultaneous Bayesian-frequentist model-averaged hypothesis testing using Doublethink

Arning, N.; Fryer, H. R.; Wilson, D. J.

2024-01-02 epidemiology
10.1101/2024.01.01.24300687 medRxiv
Show abstract

Big data approaches to discovering non-genetic risk factors have lagged behind genome-wide association studies that routinely uncover novel genetic risk factors for diverse diseases. Instead, epidemiology typically focuses on candidate risk factors. Since modern biobanks contain thousands of potential risk factors, candidate approaches may introduce bias, inadequately control for multiple testing, and overlook important signals. Doublethink, a novel model-averaged hypothesis testing approach, offers a solution that simultaneously controls the Bayesian false discovery rate (FDR) and frequentist familywise error rate (FWER) while accounting for uncertainty in variable selection. Here we investigate direct risk factors for COVID-19 hospitalization from among 1,912 variables in 201,917 UK Biobank participants by implementing a Doublethink-based exposome-wide association study using Markov Chain Monte Carlo. Focusing on the 2020 outbreak, we find nine individual variables and six groups of variables exposome-wide significant at 9% FDR and 0.05% FWER. We identify significant direct effects among relatively overlooked risk factors including psychiatric disorders, dementia and prior infection, which we evaluate in relation to studies of other populations. We detect significant direct effects among some commonly reported risk factors like age, sex and obesity, but not others like diabetes, cardiovascular disease, hypertension, which may be mediated instead through variables representing general comorbidity. Doublethink produces interchangeable posterior odds and p-values for individual variables and arbitrary groups, facilitating flexible and powerful post-hoc hypothesis testing. We discuss the potential for impact and limitations of joint Bayesian-frequentist hypothesis testing, including the benefits of an agnostic exposome-wide approach to discovery. SignificanceUnderstanding what causes disease is key to improving its treatment and prevention. Large health studies like UK Biobank measure thousands of possible causes of disease. Traditionally, scientists have studied possible causes (like smoking or exercise) one-at-a-time, in depth. For greater perspective, we could study them altogether to test which have any effect. We recently introduced Doublethink, which combines the advantages of two major statistical approaches to testing. Here we use Doublethink to test 1,912 possible causes of COVID-19 hospitalization in UK Biobank. We found strong evidence for relatively overlooked causes: psychiatric conditions, dementia and previous infections. Findings from other health studies support these causes, highlighting the need to re-evaluate them and showing how our approach can reveal valuable insights.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.