Back

Linking EMS and In-Hospital Stroke Records: Impact of Deterministic vs. Probabilistic Methods on Selection Bias

McLouth, C. J.; Goldstein, L. B.

2025-08-02 health systems and quality improvement
10.1101/2025.08.01.25332630 medRxiv
Show abstract

BackgroundLinking emergency medical services (EMS) and hospital stroke registry data is crucial for evaluating stroke systems of care, but the impact of different linkage methods on selection bias remains unclear. This study compared deterministic and probabilistic linkage approaches and assessed their effects on sample representativeness and analytical conclusions MethodsIn this cross-sectional study we analyzed 13,567 stroke patients transported by EMS to 40 Kentucky hospitals participating in Get With The Guidelines - Stroke (2021-2023). Records were linked using deterministic and probabilistic methods. We compared match rates, assessed sample representativeness, and evaluated the impact of selection bias using inverse probability weighting. ResultsDeterministic and probabilistic methods achieved match rates of 73.0% and 78.7%, respectively. Both methods produced similar representative samples, with modest differences between matched and unmatched cases primarily in race and admission year. Accounting for selection bias had minimal impact on the estimated associations between EMS stroke recognition and outcomes (percent change in adjusted odds ratios < 1%). ConclusionsWhile probabilistic linkage yielded modestly higher match rates, both methods produced comparable results with minimal selection bias. When working with high-quality data with low missingness, deterministic linkage may be sufficient for many analyses, though sensitivity analyses remain important for assessing potential bias.

Published in Neurology Open Access · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.