A Clinical Phenotyping Algorithm to Identify Cases of Chronic Obstructive Pulmonary Disease in Electronic Health Records
Martucci, V. L.; Liu, N.; Kerchberger, V. E.; Osterman, T. J.; Torstenson, E.; Richmond, B.; Aldrich, M.
Show abstract
RationaleChronic obstructive pulmonary disease (COPD) is a leading cause of mortality in the United States. Electronic health records provide large-scale healthcare data for clinical research, but have been underutilized in COPD research due to challenges identifying these individuals, especially in the absence of pulmonary function testing data.\n\nObjectivesTo develop an algorithm to electronically phenotype individuals with COPD at a large tertiary care center.\n\nMethodsWe identified individuals over 45 years of age at last clinic visit within Vanderbilt University Medical Center electronic health records. We tested phenotyping algorithms using combinations of both structured and unstructured text and examined the clinical characteristics of the resulting case sets.\n\nMeasurement and Main ResultsA simple algorithm consisting of 3 International Classification of Disease codes for COPD achieved a sensitivity of 97.6%, a specificity of 76.0%, a positive predictive value of 57.1%, and a negative predictive value of 99.0%. A more complex algorithm consisting of both billing codes and a mention of oxygen on the problem list that achieved a positive predictive value of 86.5%. However, the association of known risk factors with chronic obstructive pulmonary disease was consistent in both algorithm sets, suggesting a simple code-only algorithm may suffice for many research applications.\n\nConclusionsSimple code-only phenotyping algorithms for chronic obstructive pulmonary disease can identify case populations with epidemiologic and genetic profiles consistent with published literature. Implementation of this phenotyping algorithm will expand opportunities for clinical research and pragmatic trials for COPD.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Spirometric classifications of COPD severity as predictive markers for clinical outcomes: the HUNT Study 94%
- Blood-Based Transcriptomic and Proteomic Biomarkers of Radiologic Emphysema 93%
- Domiciliary high-flow nasal cannula oxygen therapy for stable hypercapnic COPD patients: a prospective, multicenter, open-label, randomized controlled trial 93%
Similar papers in this journal
- Prospective Analysis of urINe LAM to Eliminate NTM Sputum Screening (PAINLESS) study: Rationale and trial design for testing urine lipoarabinomannan as a marker of NTM lung infection in cystic fibrosis 92%
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 92%
- Clinical risk factors and blood protein biomarkers of 10-year pneumonia risk 91%
Similar papers in this journal
- Sotatercept Improves Small Airway Disease and Hyperinflation in Patients with Pulmonary Hypertension 90%
- Inhaled corticosteroids downregulate SARS-CoV-2-related gene expression in COPD: results from a RCT 90%
- Early initiation of corticosteroids in patients hospitalized with COVID-19 not requiring intensive respiratory support: cohort study 90%
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 90%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 90%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 89%
Similar papers in this journal
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 90%
- Clinical Study Applying Machine Learning to Detect a Rare Disease: Results and Lessons Learned 89%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.