Back

Sampling bias in large healthcare claims databases

Dahlen, A.; Charu, V.

2022-10-07 bioinformatics
10.1101/2022.10.03.510721 bioRxiv
Show abstract

Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum's de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.