Phenotyping using Structured Claims And Linked Electronic health records (PhenoSCALE): A semi-automated pipeline with an example of acute kidney injury
Pradhan, R.; Lii, J.; Wang, S.; Ball, R.; Desai, R.
Show abstract
ObjectivesProbabilistic phenotyping of health events has focused on unstructured or laboratory-based electronic health records (EHR), though drug surveillance still mostly relies on claims databases. Using acute kidney injury (AKI) as a case-example, we aimed to develop a semi-automated probabilistic phenotyping workflow using claims data. Materials and methodsWe defined highly sensitive "bronze" AKI events using ICD-10 codes and more specific "silver" events through two or more ICD-10 codes within 10 days of the bronze event. Data-driven feature selection identified co-occurring claims and the least absolute shrinkage and selection operator (LASSO) model predicting the silver labels was used to reduce the feature space and develop the final algorithms. These models were then validated against creatinine-based "gold" events derived from linked EHRs in a 20% hold out testing sample and their performance reported using area under the receiver operating characteristics curve (AUROC) and precision-recall curve (AUPRC). ResultsA total of 2144 features were identified based on co-occurrence with the AKI ICD codes. LASSO identified a set of 36 candidate features as most predictive of the silver labels, of which 7 were manually removed. The final phenotyping model with 29 claims-based features had an AUROC of 0.92 in the testing sample, compared to AUROC of 0.77 for a rule-based approach requiring 2 ICD codes for AKI. The AUPRC of the phenotyping model was 0.52. DiscussionPhenotyping algorithm developed using our semi-automated workflow outperformed rule-based approach for identifying AKI. ConclusionThis workflow may hold promise for broader application in phenotyping large-scale health data. Lay SummaryUnderstanding whether a patient has experienced a medical event like acute kidney injury (AKI) is not always straightforward when using health insurance claims data. Traditionally, researchers use simple rules--such as the presence of diagnostic codes--to identify such events, but these methods can miss cases or include incorrect ones. In this study, we developed a new semi-automated approach to improve the identification of AKI using only claims data. We first labeled potential AKI events using diagnostic codes and then used a statistical method to select other claims records that frequently occurred around these events. These patterns were used to build a prediction model, which we tested against lab-confirmed cases from linked medical records. Our model correctly identified AKI events with high accuracy and performed significantly better than the traditional rule-based method. This approach could help researchers and public health agencies more accurately identify health events in large datasets, especially when lab results or detailed clinical notes are not available.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 92%
- Clinical Study Applying Machine Learning to Detect a Rare Disease: Results and Lessons Learned 91%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 91%
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
Similar papers in this journal
- Evaluating the kidney disease progression using a comprehensive patient profiling algorithm: A hybrid clustering approach 95%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 94%
- Cardiovascular disease protein biomarkers are associated with kidney function: the Framingham Heart Study 93%
Similar papers in this journal
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 92%
- Identification of an ANCA-Associated Vasculitis Cohort Using Deep Learning and Electronic Health Records 92%
- Machine Learning Directed Interventions Associate with Decreased Hospitalization Rates in Hemodialysis Patients 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.