Back

One Size Does Not Fit All: A Data-Driven Framework for Personalized Comorbidity Scoring

Hwang, Y.-M.; Cui, Y.; Xu, J.; Pan, T.; Li, R.; Rice, B.; Tian, L.; Hernandez-Boussard, T.

2026-07-02 health informatics
10.64898/2026.06.30.26356967 medRxiv
Show abstract

Comorbidity indices are widely used in clinical research to summarize disease burden. However, traditional indices were developed decades ago in limited populations using fixed weights that do not reflect the diversity of patients in modern healthcare. We present the Personalized Comorbidity Score (PCS), a data-driven framework for context-dependent comorbidity scoring designed to capture patient complexity while remaining accessible for broad research adoption. PCS was developed using Epic Cosmos, a large national EHR network encompassing over 8 million adult inpatient encounters from 2015 to 2020, with comorbidities defined using AHRQ Clinical Classifications Software Refined categories. Models were developed separately across eight age-sex subgroups using LASSO-penalized Cox regression for feature selection and restricted mean survival time for score derivation. PCS is available in two versions: PCS Core, incorporating age, sex, and comorbidities, and PCS Extended, which additionally incorporates socioeconomic and geographic variables. PCS Core and PCS Extended achieved AUROCs of 0.812 and 0.813 for one-year mortality, outperforming traditional indices (AUROC, 0.714-0.730). PCS demonstrated consistently lower subgroup calibration error across demographic and socioeconomic groups without including race or ethnicity as model features. PCS was further evaluated in two complementary external EHR datasets (Stanford Health Care and MIMIC-IV) with distinct patient populations and data structures, where it consistently outperformed traditional indices. Open-source R and Python packages are provided to support broad adoption. PCS provides an updatable framework for comorbidity measurement that is accurate, context-dependent, and designed to evolve alongside clinical practice.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.