Augmenting Fact and Date of Death in Electronic Health Records using Internet Media Sources: A Validation Study from Two Large Healthcare Systems
LeNoue-Newton, M.; Al-Garadi, M. A.; Ngan, K.; Pillai, H. S.; Reeves, R. M.; Park, D.; Westerman, D. M.; Hernandez-Munoz, J. J.; Wang, X.; Kuzucan, A.; Wang, S.; Lin, K. J.; Fuller, C.; McPheeters, M.; Matheny, M. E.; Desai, R. J.
Show abstract
ObjectiveTo evaluate the validity of death ascertainment from publicly available internet media (IM) sources by benchmarking against state and Federal vital statics data for patients in two large healthcare systems from the US. MethodsWe extracted names and dates of birth and death from publicly available data--including obituaries and memorial websites--using previously developed natural language processing models. These data were probabilistically matched to electronic health records (EHRs) from Mass General Brigham (MGB) and Vanderbilt University Medical Center (VUMC) on first name, last name, and date of birth. Using reference standards from state vital statistics databases from MA, CT, and VT for MGB and the National Death Index (NDI) for VUMC patients, we reported positive predicted values (PPV) considering cases where dates of death from IM sources were within 7 days of the reference standard to be true positives. We also reported sensitivity of deaths ascertained from IM sources. ResultsWhen probabilistically matching 8.1 million deaths extracted from public data to 78,848 deaths observed in the reference standards across two sites, 30,607 (38.8%) matched exactly. A PPV of 98.2% for MGB and 98.9% for VUMC was observed for exact matches, while <6% for non-exact matches. Considering only the exact matches, IM sources led to an improvement in sensitivity of death capture by 24% in MGB and 18% in VUMC, compared to using EHRs alone for death ascertainment. ConclusionsUsing public information to augment mortality data increased capture of death meaningfully over reliance on EHR records alone.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 90%
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 90%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 90%
Similar papers in this journal
- Predicting the need for escalation of care or death from repeated daily clinical observations and laboratory results in patients with SARS-CoV-2 during 2020: a retrospective population-based cohort study from the United Kingdom 92%
- Missing data and missed infections: Investigating racial and ethnic disparities in SARS-CoV-2 testing and infection rates in Holyoke, Massachusetts 88%
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 92%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 92%
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 91%
Similar papers in this journal
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Systematic identification of rare disease patients in electronic health records enables evaluation of clinical outcomes 92%
- An Observational Study of COVID-19 from A Large Healthcare System in Northern New Jersey: Diagnosis, Clinical Characteristics, and Outcomes 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.