The EHR Density Index: A new method to control for EHR data inconsistency across patients
Bhatia, A.; Lash, S.; McIntee, T.; Pfaff, E.
Show abstract
Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018 - 2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing the quality of clinical and administrative data extracted from hospitals: The General Medicine Inpatient Initiative (GEMINI) experience 91%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 91%
- medExtractR: A medication extraction algorithm for electronic health records using the R programming language 91%
Similar papers in this journal
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 91%
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 91%
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 91%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 91%
Similar papers in this journal
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
- Aggregating Multiple Real-World Data Sources using a Patient-Centered Health Data Sharing Platform: an 8-week Cohort Study among Patients Undergoing Bariatric Surgery or Catheter Ablation of Atrial Fibrillation 90%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.