Bias from small-count suppression in county-level cancer disparity estimates: a calibrated simulation study
gahan, k.
Show abstract
Abstract Background. Area-level cancer disparities are routinely estimated from public county data in which rates based on small counts (fewer than 16 cases or deaths) are suppressed. Analysts typically drop suppressed counties (complete-case analysis). Because suppression depends on case counts tied to population size and demographic composition, this missingness may be informative, but its effect on the disparity estimate has not, to our knowledge, been quantified. Methods. In a cross-sectional ecological study of 3,143 U.S. counties (analytic sample 3,018 with computable exposure) using one frozen public release of NCI State Cancer Profiles incidence and mortality data and ACS 2018-2022 5-year data, we estimated the most- versus least-deprived ICE(race+income) quintile rate ratio (RR) and rate difference for female breast, stomach, and cervix cancers under four suppression-handling methods: complete-case, available-case, bounding, and model-based small-area estimation. We characterized which counties were erased, and, following the ADEMP framework, ran a Monte Carlo simulation (1,000 replicates per cell; Monte Carlo standard error of bias approximately 0.0025) calibrated to the release to measure bias against a known truth. Analyses were pre-registered. Results. The suppressed fraction rose with rarity: 7.4% of counties for breast, 61.3% for stomach, and 75.7% for cervix incidence. Suppression was concentrated in the most-deprived quintile (cervix, 81.8% suppressed vs 63.8% least-deprived) and overwhelmingly removed rural rather than minority residents (cervix: 81% of the rural but 9% of the minority population erased). For breast (little suppression) the RR was 0.87 (95% CI 0.85-0.89) and identical across methods; for cervix incidence the complete-case RR (1.56) exceeded the model-based estimate (1.50), and for cervix mortality (91% suppressed) complete-case (1.86) exceeded model-based (1.56) by 16% with a wide bounding interval (1.88-2.62). In calibrated simulation, population-weighted complete-case bias was small (less than 2%) at the observed deprivation-county-size correlation and grew with rarity, threshold, and unweighted aggregation; its direction was conditional, becoming positive (over-estimation) as deprived counties became smaller. Conclusions. Complete-case handling of suppressed counties over-estimates rare-cancer area disparities relative to methods that retain them, while silently erasing most of the rural and most-deprived communities the estimate is meant to represent. The effect is negligible for common cancers and grows with rarity. Public-data disparity analyses should report the suppressed fraction and use bounded or model-based estimates by default. Keywords: cancer disparities; small-count suppression; Index of Concentration at the Extremes; informative missingness; small-area estimation; rural health.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- How Timing of Stay-at-home Orders and Mobility Reductions Impacted First-Wave COVID-19 Deaths in US Counties 93%
- Social inequalities in COVID-19 deaths by area-level income: patterns over time and the mediating role of vaccination in a population of 11.2 million people in Ontario, Canada 91%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 91%
Similar papers in this journal
- Disentangling the relationship between cancer mortality and COVID-19 92%
- The impact of COVID-19 on population cancer screening programs in Australia: modelled evaluations for breast, bowel and cervical cancer 91%
- Health impacts of COVID-19 disruptions to primary cervical screening by time since last screen: A model-based analysis for current and future disruptions 91%
Similar papers in this journal
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 92%
- COVerAGE-DB: A database of age-structured COVID-19 cases and deaths 90%
- Potential Test-Negative Design Study Bias in Outbreak Settings: Application to Ebola vaccination in Democratic Republic of Congo 90%
Similar papers in this journal
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 91%
- A framework to model global, regional, and national estimates of intimate partner violence 90%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 89%
Similar papers in this journal
- Deprivation and Segregation in Ovarian cancer survival among African American Women: a mediated analysis 91%
- Racial/Ethnic Disparities in the Observed COVID-19 Case Fatality Rate Among the U.S. Population 90%
- Spatially refined time-varying reproduction numbers of SARS-CoV-2 in Arkansas and Kentucky and their relationship to population size and public health policy, March – November, 2020 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.