Sampling bias in large healthcare claims databases
Dahlen, A.; Charu, V.
Show abstract
Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum's de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Measure what matters: counts of hospitalized patients are a better metric for health system capacity planning for a reopening 91%
- Assessing the quality of clinical and administrative data extracted from hospitals: The General Medicine Inpatient Initiative (GEMINI) experience 90%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 89%
Similar papers in this journal
- Characterization of Long COVID among U.S. Medicare Beneficiaries using Claims Data 91%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 90%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 89%
Similar papers in this journal
- An Observational Study of COVID-19 from A Large Healthcare System in Northern New Jersey: Diagnosis, Clinical Characteristics, and Outcomes 90%
- Association of Chronic Acid Suppression and Social Determinants of Health with COVID-19 Infection 90%
- Racial and ethnic disparities for SARS-CoV-2 positivity in the United States: a generalizing pandemic 90%
Similar papers in this journal
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 90%
- COVID-19 collateral: Indirect acute effects of the pandemic on physical and mental health in the UK 88%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.