Clinical encounter heterogeneity and methods for resolving in networked EHR data: A study from N3C and RECOVER programs
Leese, P. J.; Anand, A.; Girvin, A.; Bennett, T.; Hajagos, J.; Patel, S.; Yoo, J.; Pfaff, E.; Moffitt, R.
Show abstract
OBJECTIVEClinical encounter data are heterogeneous and vary greatly from institution to institution. These problems of variance affect interpretability and usability of clinical encounter data for analysis. These problems are magnified when multi-site electronic health record data are networked together. This paper presents a novel, generalizable method for resolving encounter heterogeneity for analysis by combining related atomic encounters into composite macrovisits. MATERIALS AND METHODSEncounters were composed of data from 75 partner sites harmonized to a common data model as part of the NIH Researching COVID to Enhance Recovery Initiative, a project of the National Covid Cohort Collaborative. Summary statistics were computed for overall and site-level data to assess issues and identify modifications. Two algorithms were developed to refine atomic encounters into cleaner, analyzable longitudinal clinical visits. RESULTSAtomic inpatient encounters data were found to be widely disparate between sites in terms of length-of-stay and numbers of OMOP CDM measurements per encounter. After aggregating encounters to macrovisits, length-of-stay (LOS) and measurement variance decreased. A subsequent algorithm to identify hospitalized macrovisits further reduced data variability. DISCUSSIONEncounters are a complex and heterogeneous component of EHR data and native data issues are not addressed by existing methods. These types of complex and poorly studied issues contribute to the difficulty of deriving value from EHR data, and these types of foundational, large-scale explorations and developments are necessary to realize the full potential of modern real world data. CONCLUSIONThis paper presents method developments to manipulate and resolve EHR encounter data issues in a generalizable way as a foundation for future research and analysis.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Real World Performance of the 21st Century Cures Act Population Level Application Programming Interface 96%
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 93%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 93%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 96%
- Signal from the Noise: A Mixed Methods Process Mining Approach to Evaluate Care Pathways. 94%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 96%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.