Using fragmented data to characterize community healthcare utilization
McCready, T.; Thorpe, L.; Roy, B.; Renson, A.
Show abstract
Community-level estimates of healthcare utilization are essential for identifying inequities, allocating resources, and evaluating place-based interventions. However, in the United States, no single data source adequately captures healthcare utilization within geographically defined populations. Population-based surveys often lack sufficient geographic resolution, insurance claims represent only covered populations, and electronic health records are limited to care delivered within participating health systems. Increasingly, researchers combine these fragmented data sources, yet limited guidance exists for conducting valid population-based descriptive analyses using incomplete and overlapping data. We review the strengths and limitations of major data sources used to characterize community healthcare utilization and propose an approach for conducting population-based descriptive analyses using fragmented data. Rather than focusing on the limitations of individual data sources, our approach begins by explicitly defining the target population and the ideal observational study that would answer the research question. Available data sources are then conceptualized as incomplete or imperfect realizations of that ideal, providing a structured approach to (a) identifying sources of selection bias, missingness, and measurement error, (b) articulating required assumptions, and (c) selecting appropriate analytic strategies. We illustrate our approach using colorectal cancer screening utilization among adults residing in Brooklyn, New York during 2022. By shifting attention from individual data sources to the target community and the assumptions required for valid inference, this approach provides a practical approach for strengthening descriptive analyses of community healthcare utilization and informing place-based public health research, policy, and practice.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- County-Level Estimates of Excess Mortality associated with COVID-19 in the United States 91%
- Life Expectancy and Voting Patterns in the 2020 U.S. Presidential Election 90%
- Financial Hardship and Social Assistance as Determinants of Mental Health and Food and Housing Insecurity During the COVID-19 Pandemic 90%
Similar papers in this journal
- Measure what matters: counts of hospitalized patients are a better metric for health system capacity planning for a reopening 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 92%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 92%
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 94%
- A scoping review of fair machine learning techniques when using real-world data 91%
- Demonstrating the Consequences of Learning Missingness Patterns in Early Warning Systems for Preventative Health Care: A Novel Simulation and Solution 91%
Similar papers in this journal
- Bias-adjusted predictions of county-level vaccination coverage from the COVID-19 Trends and Impact Survey 92%
- A novel decision modeling framework for health policy analyses when outcomes are influenced by social and disease processes 91%
- Forecasting local surges in COVID-19 hospitalizations through adaptive decision tree classifiers 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.