Multi-stage reweighting to correct for participation bias in a nationwide biobank with nested recruitment
Traeholt, J.; Didriksen, M.; Helenius, D.; Christoffersen, L. A. N.; Dinh, K. M.; Dowsett, J.; Mikkelsen, C.; Hindhede, L.; Quinn, L. J. E.; Bruun, M. T.; Aagaard, B.; Hansen, T. F.; Hjalgrim, H.; Rostgaard, K.; Sorensen, E.; Erikstrup, C.; Pedersen, O. B. V.; Hansen, T.; Schork, A. J.; Markussen, B.; Ostrowski, S. R.
Show abstract
Selective participation in biobanks often compromises inference to the general population, particularly when selection occurs across multiple stages, whether at recruitment or during subsequent participation. Inverse probability (IP) weighting can reduce systematic differences using suitable external benchmarks, but most applications assume a single selection process. Here, we present a multi-stage IP-weighting framework and apply it to the Danish Blood Donor Study (DBDS), a nationwide biobank embedded in Denmark's blood-donation infrastructure. Using national registers, we estimated year-specific probabilities of (i) donation activity and (ii) DBDS enrolment conditional on donation activity, yielding two-stage inclusion weights for 169,893 participants. These weights reduced inclusion-associated imbalance across the 52 auxiliary variables in the probability models by 97.6% (median) and, despite strong health selection under donation-based recruitment, reduced relative-prevalence discrepancies across held-out prescription phenotypes by 69.7% (median). The effective sample size after weighting was 30,627 (18.0% of 169,893). Combining the inclusion weights with questionnaire-specific response weights across five DBDS questionnaires (>500 questions) produced the largest changes from unweighted to weighted responses for health behaviours and symptom severity, including tobacco and alcohol consumption, menstrual-pain severity, restless-legs severity, nocturia, sleep disturbance, and fatigue. These findings support multi-stage IP-weighting to improve population alignment in biobanks with staged selection.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk 94%
- Within-family studies for Mendelian randomization: avoiding dynastic, assortative mating, and population stratification biases 93%
- Genetic predictors of participation in optional components of UK Biobank 92%
Similar papers in this journal
- Protection of previous SARS-CoV-2 infection is similar to that of BNT162b2 vaccine protection: A three-month nationwide experience from Israel 90%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 89%
- A system for phenotype harmonization in the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program 89%
Similar papers in this journal
- Polygenic Risk Score Improves the Accuracy of a Clinical Risk Score for Coronary Artery Disease 92%
- Meat consumption and risk of 25 common conditions: outcome-wide analyses in 475,000 men and women in the UK Biobank study 90%
- Effectiveness and efficiency of immunisation strategies to prevent RSV among infants and older adults in Germany: a modelling study 90%
Similar papers in this journal
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 93%
- Association between household composition and severe COVID-19 outcomes in older people by ethnicity: an observational cohort study using the OpenSAFELY platform 91%
- The Causal Effects of Health Conditions and Risk Factors on Social and Socioeconomic Outcomes: Mendelian Randomization in UK Biobank 91%
Similar papers in this journal
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 93%
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 92%
- No statistical evidence for an effect of CCR5-Δ32 on lifespan in the UK Biobank cohort 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.