Who seeks care, and what gets measured? Understanding the distinct mechanisms behind visit and observation processes in multi-center electronic health records
Yang, C.-H.; Salvatore, M.; Lu, H.; Zhu, Z.; Tennant, P.; Shi, X.; Ohno-Machado, L.; Khera, R.; Gross, C.; Li, F.; Mukherjee, B.
Show abstract
Electronic health record (EHR)-linked cohorts support association, prediction, and causal studies using longitudinally measured markers of health. However, a lab biomarker measurement is recorded only when a patient first has a medical encounter (visit process) and, a clinician orders the corresponding test and the patient follows through (observation process). These two stages may induce informative presence (IP) and informative observation (IO), respectively. Yet their drivers remain largely uncharacterized, despite evidence that understanding this recording mechanism is essential for selecting appropriate strategies for downstream analysis that treat these markers as longitudinally measured outcomes. We characterize this two-stage recording hierarchy using a stochastic recurrent-event model for the outpatient visit process and a visit-process-weighted generalized estimating equation model for biomarker recording conditional on an outpatient visit. We characterize descriptors of both processes in three EHR-linked cohorts in the US (All of Us [AoU], n=599,423; Yale New Haven Health System [YNHHS], n=319,666; Michigan Genomics Initiative [MGI], n=82,372), reporting descriptive statistics for longitudinal visits and for a panel of 68 lab biomarkers commonly measured in EHRs. We conduct detailed model-based analyses of ten biomarkers spanning multiple domains: routine monitoring, general laboratory assessment, and symptom-triggered testing. These include glucose, hemoglobin A1c [HbA1c], creatinine, hemoglobin [Hgb], white blood cell count [WBC], low-density lipoprotein [LDL] and high-density lipoprotein [HDL] cholesterol, triglycerides, C-reactive protein [CRP], and thyroid-stimulating hormone [TSH]. Across the three cohorts, the median number of outpatient visits ranged from 1.7 to 6.1 per year over a median follow-up of 4.4 to 7.2 years. Among patients with at least one recorded measurement, the median within-person proportion of visits containing a given biomarker ranged from 0.4% to 19.5%, demonstrating that more frequent visits did not necessarily translate into greater per-visit biomarker capture. In the visit-process models, chronic disease burden, and a recent history of outpatient visits were consistently associated with higher visit rates across all three cohorts whereas associations with race, ethnicity, and neighborhood-level income varied across cohorts. In per-visit observation models, the association of covariates depended on the biomarker under consideration; for example, prior cancer diagnosis was associated with more frequent measurement of blood counts but with less frequent measurement of lipids. These findings provide a deeper understanding of how to model who seeks care and what is measured as two distinct recording processes in EHR. Our empirical findings show that the descriptors of these processes vary across cohorts and biomarkers, providing guidance on how to construct these models for downstream longitudinal analyses with irregular EHR visits.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 91%
- High-throughput Phenotyping with Temporal Sequences 91%
- Measure what matters: counts of hospitalized patients are a better metric for health system capacity planning for a reopening 91%
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 92%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 91%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 91%
Similar papers in this journal
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 91%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 91%
- Demonstrating the Consequences of Learning Missingness Patterns in Early Warning Systems for Preventative Health Care: A Novel Simulation and Solution 91%
Similar papers in this journal
Similar papers in this journal
- Forecasting hospital-level COVID-19 admissions using real-time mobility data 90%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 89%
- The prevalence of SARS-CoV-2 infection and other public health outcomes during the BA.2/BA.2.12.1 surge, New York City, April-May 2022 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.