A reproducible open-source framework for defining type 1 and type 2 diabetes research cohorts in routinely collected electronic health record data
Hopkins, R.; Cardoso, P.; Güdemann, L. M.; McGovern, A. P.; Dennis, J. M.; Shields, B. M.; Young, K. G.
Show abstract
BackgroundElectronic health record data (EHR) data provide an increasingly important resource for studying people living with diabetes and their clinical outcomes, but robustly coding reproducible datasets is challenging. We aimed to develop a standardised data-processing framework for defining cohorts of people with type 1 and type 2 diabetes using EHR data. MethodsWe initially provide a standardised, generalisable procedure to develop clinically reviewed code lists to robustly define variables in EHR data. Using UK population-based data from primary care linked to hospital admission records (Clinical Practice Research Datalink [CPRD]), we develop and demonstrate a data-processing pipeline applicable to raw EHR data, using clinical code lists to define a population of individuals with diabetes and defining their diabetes diagnosis dates using the earliest recorded observation of diabetes (clinical code, high HbA1c test result, or prescription for glucose lowering therapy). Using a previously validated approach, we classify diabetes type (gold standard type 1, type 2) based on insulin prescriptions, diabetes type specific clinical codes, and age at diagnosis. Finally, we demonstrate how multiple research cohorts can be defined from this diabetes population based on a specific index date, including a range of baseline features (sociodemographic and lifestyle factors, biomarkers, comorbidities, medications) and key outcomes relevant to the research question. ResultsApplication of the framework identified an incident cohort at diabetes diagnosis (type 1 diabetes (T1D): N = 10,480, mean age at diagnosis [SD] = 10.4 [4.8]; type 2 diabetes (T2D): N = 726,800, mean age at diagnosis [SD] = 60.5 [13.4]), a prevalent cohort actively registered with their GP practice on 01/02/2020 (T1D: N = 9,514, T2D: N = 559,905), and a T2D cohort initiating treatment with glucose- lowering therapies (N = 769,394 treatment initiations, considering 7 major medication classes). We publicly share our code lists and data processing code, making our research as transparent and reproducible as possible (https://github.com/Exeter-Diabetes/CPRD-Cohort-scripts, https://github.com/Exeter-Diabetes/CPRD-Codelists/). ConclusionsWe have developed a flexible and reproducible framework to generate analysis-ready diabetes research cohorts in EHR data. The concepts of this framework are applicable to any EHR dataset and have been shared for use by other researchers. This approach could improve the quality and reproducibility of the diverse epidemiological and clinical diabetes studies using EHR worldwide.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The iDiabetes Platform: Enhanced Phenotyping of Patients with Diabetes for Precision Diagnosis, Prognosis and Treatment- study protocol for a cluster-randomised controlled study 96%
- Replicating a COVID-19 study in a national England database to assess the generalisability of research with regional electronic health record data 95%
- Determining the feasibility of calculating pancreatic cancer risk scores for people with new-onset diabetes in primary care (DEFEND PRIME): study protocol 95%
Similar papers in this journal
- Replication and cross-validation of T2D subtypes based on clinical variables: an IMI-RHAPSODY study 93%
- Development and validation of a Trans-Ancestry polygenic risk score for Type 1 Diabetes 93%
- Young onset diabetes in Asian Indians is associated with lower measured and genetically determined beta-cell function: an INSPIRED study 92%
Similar papers in this journal
- Characterisation of type 2 diabetes subgroups and their association with ethnicity and clinical outcomes: a UK real-world data study using the East London Database 94%
- Development and validation of a modified Cambridge Multimorbidity Score for use with internationally recognized electronic health record clinical terms (SNOMED CT) 92%
- Weight trends amongst adults with diabetes or hypertension during the COVID-19 pandemic: an observational study using OpenSAFELY 91%
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 94%
- Illustrating Potential Effects of Alternate Control Populations on Real-World Evidence-based Statistical Analyses 92%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.