Algorithmic Ascertainment of Cause of Death from Longitudinal Real-World Medical Claims Data: Development and Validation
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Strengthening Policy Coding Methodologies to Improve COVID-19 Disease Modeling and Policy Responses: A Proposed Coding Framework and Recommendations 90%
- Towards reduction in bias in epidemic curves due to outcome misclassification through Bayesian analysis of time-series of laboratory test results: Case study of COVID-19 in Alberta, Canada and Philadelphia, USA 89%
- Comparing methods to predict baseline mortality for excess mortality calculations 88%
Similar papers in this journal
Similar papers in this journal
- Post-discharge Acute Care and Outcomes in the Era of Readmission Reduction: A National Retrospective Cohort Study of Medicare Beneficiaries in the United States 90%
- Risk of hospitalisation with coronavirus disease 2019 in healthcare workers and their households:a nationwide linkage cohort study 89%
- Incidence, clinical outcomes, and transmission dynamics of hospitalized 2019 coronavirus disease among 9,596,321 individuals residing in California and Washington, United States: a prospective cohort study 88%
Similar papers in this journal
- Estimating and Testing an Index of Bias Attributable to Composite Outcomes in Comparative Studies 91%
- Protocol for an observational study evaluating new approaches to modelling diagnostic information from large administrative hospital datasets 90%
- Performance of ICD-10-based injury severity scores in pediatric trauma patients using the ICD-AIS map and survival rate ratios 90%
Similar papers in this journal
- Estimating excess mortality in people with cancer and multimorbidity in the COVID-19 emergency 90%
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 90%
- Development of a resilience assessment tool for cardiac care pathways in Europe: A mixed-methods study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.