Benchmarking commercial healthcare claims data
Dahlen, A.; Deng, Y.; Charu, V.
Show abstract
ImportanceCommercial healthcare claims datasets represent a sample of the US population that is biased along socioeconomic/demographic lines; depending on the target population of interest, results derived from these datasets may not generalize. Rigorous comparisons of claims-derived results to ground-truth data that quantify this bias are lacking. Objectives(1) To quantify the extent and variation of the bias associated with commercial healthcare claims data with respect to different target populations; (2) To evaluate how socioeconomic/demographic factors may explain the magnitude of the bias. DesignThis is a retrospective observational study. Healthcare claims data come from the Merative MarketScan(R) Commercial Database; reference data for comparison come from the State Inpatient Databases (SID) and the US Census. We considered three target populations, aged 18-64 years: (1) all Americans; (2) Americans with health insurance; (3) Americans with commercial health insurance. ParticipantsWe analyzed inpatient discharge records of patients aged 18-64 years, occurring between 01/01/2019 to 12/31/2019 in five states: California, Iowa, Maryland, Massachusetts, and New Jersey. OutcomesWe estimated rates of the 250 most common inpatient procedures, using claims data and using reference data for each target population, and we compared the two estimates. ResultsThe average rate of inpatient discharges per 100 person-years was 5.39 in the claims data (95% CI: [5.37, 5.40]) and 7.003 (95% CI: [7.002, 7.004]) in the reference data for all Americans, corresponding to a 23.1% underestimate from claims. We found large variation in the extent of relative bias across inpatient procedures, including 22.8% of procedures that were underestimated by more than a factor of 2. There was a significant relationship between socioeconomic/demographic factors and the magnitude of bias: procedures that disproportionately occur in disadvantaged neighborhoods were more underestimated in claims data (R2 = 51.6%, p < 0.001). When the target population was restricted to commercially insured Americans, the bias decreased substantially (3.2% of procedures were biased by more than factor of 2), but some variation across procedures remained. Conclusions and relevanceNaive use of healthcare claims data to derive estimates for the underlying US population can be severely biased. The extent of bias is at least partially explained by neighborhood-level socioeconomic factors.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benzodiazepine Initiation Effect on Mortality Among Medicare Beneficiaries Post Acute Ischemic Stroke 91%
- Using quantitative bias analysis to adjust for misclassification of COVID-19 outcomes: An applied example of inhaled corticosteroids and COVID-19 outcomes 91%
- Characterization of Long COVID among U.S. Medicare Beneficiaries using Claims Data 91%
Similar papers in this journal
- Augmenting Fact and Date of Death in Electronic Health Records using Internet Media Sources: A Validation Study from Two Large Healthcare Systems 91%
- The US Midlife Mortality Crisis Continues: Excess Cause-Specific Mortality During 2020 91%
- How Timing of Stay-at-home Orders and Mobility Reductions Impacted First-Wave COVID-19 Deaths in US Counties 91%
Similar papers in this journal
- Association of Chronic Acid Suppression and Social Determinants of Health with COVID-19 Infection 93%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 91%
- Comparison of COVID-19 outcomes among shielded and non-shielded populations: A general population cohort study of 1.3 million 91%
Similar papers in this journal
- Measuring the missing: greater racial and ethnic disparities in COVID-19 burden after accounting for missing race/ethnicity data 92%
- Evaluating the impact of keeping indoor dining closed on COVID-19 rates among large US cities: a quasi-experimental design 91%
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.