Investigating Primary Care Indications to Improve the Quality of Electronic Health Record Data in Target Trial Emulation for Dementia
Sunog, M.; Magdamo, C.; Charpignon, M.-L.; Albers, M. W.
Show abstract
Missing data, inaccuracies in medication lists, and recording delays in electronic health records (EHR) are major limitations for target trial emulation (TTE), which uses EHR data to retrospectively emulate a clinical trial. EHR-based TTE relies on recorded data that proxy actual drug exposures and outcomes. While prior work has proposed various methods to improve EHR data quality, here we investigate the underutilized consideration that encounters with a primary care provider (PCP) may result in more accurate data in the EHR. Patients with a PCP within the EHR network being studied tend to have more encounters overall and a greater proportion of the types of encounters that yield comprehensive and up-to-date records. By contrasting data for patients with and without a PCP in the considered EHR network, we demonstrate how PCP status affects EHR data quality. Through a case study, we then empirically examine the impact on TTE of including a PCP status feature either in the propensity score and outcome models or as an eligibility criterion for cohort selection, versus ignoring it. Specifically, we compare the estimated effects of two first-line antidiabetic drug classes on the onset of Alzheimers Disease and Related Dementias. We find that the estimated treatment effect is sensitive to the consideration of PCP status, particularly when used as an eligibility criterion. Our work suggests that further researching the role of PCP status may improve the design of pragmatic trials. Data and Code AvailabilityThe study uses EHR data from the Research Patient Data Registry (Nalichowski et al., 2007), social vulnerability index (SVI) data from the Agency for Toxic Substances and Disease Registry (https://www.atsdr.cdc.gov/placeandhealth/svi), and Massachusetts death records from the Registry of Vital Records and Statistics. Because the data contain patient information, they cannot be made available. Institutional Review Board (IRB)This research was performed under MGB IRB approval (protocol 2023P000604).
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 93%
Similar papers in this journal
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 92%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 92%
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 91%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 91%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 91%
Similar papers in this journal
- CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization 93%
- Patterns of rates of mortality in the Clinical Practice Research Datalink 93%
- Effect of common maintenance drugs on the risk and severity of COVID-19 in elderly patients 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.