Back

Increasing Representativeness in the All of Us Cohort Using Inverse Probability Weighting

Kambara, M. S.; Sharma, S.; Spouge, J. L.; Jordan, I. K.; Marino-Ramirez, L.

2024-10-02 genetic and genomic medicine
10.1101/2024.10.02.24314774 medRxiv
Show abstract

Large-scale population biobanks rely on volunteer participants, which may introduce biases that compromise the external validity of epidemiological studies. We characterized the volunteer participant bias for the All of Us Research Program cohort and developed a set of inverse probability (IP) weights that can be used to mitigate this bias. The All of Us cohort is older, more female, more likely to have higher education, more likely to be covered by health insurance, less White, less likely to drink or smoke, and less likely to report being healthy compared to the US population. IP weights developed via comparison of a nationally representative database reduced the observed biases for all demographic and lifestyle characteristics. Furthermore, IP weighting corrected for differences in the correlation structure of the data. For all variables we corrected for, IP weighting brought correlation coefficients and pairwise variable associations closer to the nationally representative estimates. We provide our IP weights as a community resource to increase the representativeness and external validity of the All of Us cohort.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.