One Size Does Not Fit All: A Data-Driven Framework for Personalized Comorbidity Scoring
Hwang, Y.-M.; Cui, Y.; Xu, J.; Pan, T.; Li, R.; Rice, B.; Tian, L.; Hernandez-Boussard, T.
Show abstract
Comorbidity indices are widely used in clinical research to summarize disease burden. However, traditional indices were developed decades ago in limited populations using fixed weights that do not reflect the diversity of patients in modern healthcare. We present the Personalized Comorbidity Score (PCS), a data-driven framework for context-dependent comorbidity scoring designed to capture patient complexity while remaining accessible for broad research adoption. PCS was developed using Epic Cosmos, a large national EHR network encompassing over 8 million adult inpatient encounters from 2015 to 2020, with comorbidities defined using AHRQ Clinical Classifications Software Refined categories. Models were developed separately across eight age-sex subgroups using LASSO-penalized Cox regression for feature selection and restricted mean survival time for score derivation. PCS is available in two versions: PCS Core, incorporating age, sex, and comorbidities, and PCS Extended, which additionally incorporates socioeconomic and geographic variables. PCS Core and PCS Extended achieved AUROCs of 0.812 and 0.813 for one-year mortality, outperforming traditional indices (AUROC, 0.714-0.730). PCS demonstrated consistently lower subgroup calibration error across demographic and socioeconomic groups without including race or ethnicity as model features. PCS was further evaluated in two complementary external EHR datasets (Stanford Health Care and MIMIC-IV) with distinct patient populations and data structures, where it consistently outperformed traditional indices. Open-source R and Python packages are provided to support broad adoption. PCS provides an updatable framework for comorbidity measurement that is accurate, context-dependent, and designed to evolve alongside clinical practice.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 95%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 94%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 94%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 92%
Similar papers in this journal
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 93%
- Predictive performance and clinical application of COV50, a urinary proteomic biomarker in early COVID-19 infection: a cohort study 92%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 91%
Similar papers in this journal
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 92%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 90%
- Evaluation of a Machine Learning Approach Utilizing Wearable Data for Prediction of SARS-CoV-2 Infection in Healthcare Workers 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.