Back

The EHR Density Index: A new method to control for EHR data inconsistency across patients

Bhatia, A.; Lash, S.; McIntee, T.; Pfaff, E.

2026-08-06 health informatics
10.64898/2026.08.03.26359595 medRxiv
Show abstract

Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018 - 2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.