Self-report inaccuracy in the UK Biobank: Impact on inference and interplay with selective participation
Schoeler, T.; Pingault, J.-B.; Kutalik, Z.
Show abstract
While the use of short self-report measures is common practice in biobank initiatives, such phenotyping strategy is inherently prone to reporting errors. In this work, we aimed to explore challenges related to self-report errors for biobank-scale research. We derived a reporting error score (RESUM) for n=73,129 UK Biobank (UKBB) participants, capturing inconsistent self-reporting in time-invariant phenotypes across multiple measurement occasions. We then performed genome-wide association scans on RESUM, applied downstream analyses (LD Score Regression and Mendelian Randomization, MR), and compared its properties to a previously studied participation behaviour (UKBB participation propensity). The results were then used in extended analyses (simulations, inverse probability and variance weighting) to explore patterns and propose possible corrections for biases induced by reporting error and/or selective participation. Finally, to assess the impact of reporting error on SNP effects and trait heritability, we improved phenotype resolution for 15 self-report measures and inspected the changes in genomic findings. Reporting error was present in the UKBB across all 33 assessed, time-invariant, measures, with repeatability levels as low as 11% (e.g., inconsistent recall of childhood sunburns). We found that reporting error was not independent from UKBB participation, evidenced by their negative genetic correlation (rg = -0.90), their shared causes (e.g., education, income, intelligence; assessed in MR) and the loss in self-report accuracy following participation bias correction. Depending on where reporting error occurred in the analytical pipeline, its impact ranged from reduced power (e.g., for gene-discovery) to biased effect estimates (e.g., if present in the exposure variable) and attenuation of genome-wide quantities (e.g., 20% relative h2-attenuation for self-reported childhood height). Our findings highlight that both self-report accuracy and selective participation are competing biases and sources of poor reproducibility for biobank-scale research. Implementation of approaches that aim to enhance phenotype resolution while ensuring sample representativeness are therefore essential when working with biobank data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Within-family studies for Mendelian randomization: avoiding dynastic, assortative mating, and population stratification biases 96%
- Genetic predictors of participation in optional components of UK Biobank 96%
- A novel method for an unbiased estimate of cross-ancestry genetic correlation using individual-level data 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 94%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 94%
- Investigating a possible causal relationship between maternal serum urate concentrations and offspring birthweight: A Mendelian randomization study 93%
Similar papers in this journal
- Correction for participation bias in the UK Biobank reveals non-negligible impact on genetic associations and downstream analyses 99%
- Resource Profile and User Guide of the Polygenic Index Repository 97%
- Rank concordance of polygenic indices: Implications for personalised intervention and gene-environment interplay 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.