Causal Inference via Electronic Health Records in the National Clinical Cohort Collaborative: Challenges and Solutions in Long COVID Research
Butzin-Dozier, Z.; Ji, Y.; Wang, L.-C.; Anzalone, A. J.; Hurwitz, E.; Patel, R. C.; van der Laan, M.; Colford, J. M.; Hubbard, A. E.; on behalf of the N3C Consortium,
Show abstract
Observational analyses of electronic health record (EHR) data using databases such as the National Clinical Cohort Collaborative include unique challenges for researchers seeking causal inferences, particularly when evaluating subjectively-defined outcomes like Long COVID. We explore several challenges and describe potential solutions. 1. Lack of true negatives: Many diagnoses and conditions either have a positive indicator or a missing status, requiring investigators to carefully consider which patients are likely negative for this condition. 2. Differential monitoring: EHR data include nonrandom missingness driven by patients engaging with the healthcare system at different rates, which is often related to both the exposure and outcome of interest. 3. Bias: EHR data sources face many biases, but are particularly vulnerable to informative missingness, differential monitoring, and model misspecification. 4. Large sample size: High precision (i.e., narrow confidence intervals) paired with potential bias leads to a high risk of incorrectly rejecting the null hypothesis. 5. Defining index time: It is important that investigators deliberately define index time (i.e., t0, baseline) to ensure that they only adjust for baseline confounders and do not adjust for (or condition on) factors that are affected by the exposure of interest (i.e., colliders or mediators). 6. Parameter selection: Investigators should only select parameters that are supported by the data distribution. This manuscript provides an overview of these challenges and solutions, using both simulated data and real-world data, with the outcome of Long COVID as the running example.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 94%
- Theoretical framework for retrospective studies of the effectiveness of SARS-CoV-2 vaccines 92%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 92%
Similar papers in this journal
- Potential Biases in Test-Negative Design Studies of COVID-19 Vaccine Effectiveness Arising from the Inclusion of Asymptomatic Individuals 94%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 92%
- Estimating protection afforded by prior infection in preventing reinfection: Applying the test-negative study design 91%
Similar papers in this journal
- Bias amplification of unobserved confounding in pharmacoepidemiological studies using indication-based sampling: there is no free lunch in restricting the sample to those with a particular drug-indication 95%
- Benzodiazepine Initiation Effect on Mortality Among Medicare Beneficiaries Post Acute Ischemic Stroke 92%
- Using quantitative bias analysis to adjust for misclassification of COVID-19 outcomes: An applied example of inhaled corticosteroids and COVID-19 outcomes 92%
Similar papers in this journal
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 93%
- Toward Evaluation of Disseminated Effects of Medications for Opioid Use Disorder within Provider-Based Clusters Using Routinely-Collected Health Data 93%
- Efficient Estimation of Indirect Effects in Case-Control Studies Using a Unified Likelihood Framework 93%
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 96%
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 94%
- A scoping review of fair machine learning techniques when using real-world data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.