Back

Augmenting Fact and Date of Death in Electronic Health Records using Internet Media Sources: A Validation Study from Two Large Healthcare Systems

LeNoue-Newton, M.; Al-Garadi, M. A.; Ngan, K.; Pillai, H. S.; Reeves, R. M.; Park, D.; Westerman, D. M.; Hernandez-Munoz, J. J.; Wang, X.; Kuzucan, A.; Wang, S.; Lin, K. J.; Fuller, C.; McPheeters, M.; Matheny, M. E.; Desai, R. J.

2025-01-27 epidemiology
10.1101/2025.01.24.25321042 medRxiv
Show abstract

ObjectiveTo evaluate the validity of death ascertainment from publicly available internet media (IM) sources by benchmarking against state and Federal vital statics data for patients in two large healthcare systems from the US. MethodsWe extracted names and dates of birth and death from publicly available data--including obituaries and memorial websites--using previously developed natural language processing models. These data were probabilistically matched to electronic health records (EHRs) from Mass General Brigham (MGB) and Vanderbilt University Medical Center (VUMC) on first name, last name, and date of birth. Using reference standards from state vital statistics databases from MA, CT, and VT for MGB and the National Death Index (NDI) for VUMC patients, we reported positive predicted values (PPV) considering cases where dates of death from IM sources were within 7 days of the reference standard to be true positives. We also reported sensitivity of deaths ascertained from IM sources. ResultsWhen probabilistically matching 8.1 million deaths extracted from public data to 78,848 deaths observed in the reference standards across two sites, 30,607 (38.8%) matched exactly. A PPV of 98.2% for MGB and 98.9% for VUMC was observed for exact matches, while <6% for non-exact matches. Considering only the exact matches, IM sources led to an improvement in sensitivity of death capture by 24% in MGB and 18% in VUMC, compared to using EHRs alone for death ascertainment. ConclusionsUsing public information to augment mortality data increased capture of death meaningfully over reliance on EHR records alone.

Published in American Journal of Epidemiology (predicted rank #3) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.